Methodology

PureGameStats publishes proprietary estimates for sales, wishlists, and pre-launch momentum. This page describes the inputs each model uses, how we validate its output, and where each estimate should be trusted versus taken as a directional signal.

Last methodology review: 3 September 2026

What we do not do: publish coefficients, weights, or the specific formulas that generate each number. Those are the parts of the model that took time to tune and remain proprietary. Everything else is documented below.

Sales estimate

Model
PGS Sales Model, current generation.
Data inputs
Steam review totals and positive-to-negative ratio, genre bands, tag dampeners for known outlier categories, launch cadence and days-since-release, regional review-language mix (used to detect regional-market titles that under-report reviews), Early Access graduation signal, viral-peak anchoring for mega-launches.
Validation
Continuously cross-referenced against developer-reported sales figures. Every new dev-published number is a calibration anchor.
Confidence
High for released commercial titles more than 30 days post-launch.
Medium for pre-release estimates and titles inside their launch window.
Lower for free-to-play viral hits where community gifting and platform-cross-play distort review counts.

Sales estimates should be treated as modelled ranges rather than audited sales figures. They are designed primarily for comparative analysis across titles and for estimating the likely order of magnitude of an individual game's performance.

Wishlist estimate

Model
A segmented model combining Steam wishlist rank and follower behaviour.
Data inputs
Steam public wishlist rank when disclosed by Valve, follower count, first-tag genre multiplier, Early Access flag, follower conversion rate for rank-stable titles, segment-straddle detection for fast movers.
Validation
Cross-checked against developer-disclosed wishlist totals whenever those become public.
Confidence
High for titles inside Steam's disclosed rank window.
Medium for long-tail ranks and titles with a stable rank but rapid follower growth.
Advisory for unranked titles with fewer than a few hundred followers.

Wishlist estimates for event-window debuts (Gamescom, SGF, Next Fest, TGA) tend to run below the eventual total because these games arrive via discovery outside a normal follower base. We do not retune the model against these events since it would drag accuracy in normal weeks.

Potential Hits momentum

Model
Composite score, ranked output.
Data inputs
Follower velocity, press-coverage volume (major gaming outlets), live Twitch presence, Steam public wishlist rank, Google Trends search interest, franchise-recognition, Discord community.
Validation
Retrospective, not predictive. We track how the score at T-30 days correlates with subsequent launch-week peak concurrent users. The board self-corrects as new launches age in and out of the window.
Confidence
Advisory ranking rather than a numeric forecast. Useful for spotting which unreleased and just-released games are showing signs the other lists do not surface.

Sudden viral hits remain a structural blind spot for any pre-launch model that reads Steam-native signals. PEAK, Schedule I and similar titles can go from very limited pre-launch signal to chart-topping CCU in weeks. We accept this and do not tune the model around exceptional viral events. However, the combination of inputs can still provides early indicators of breakout behaviour beyond the normal parameters.

Review word cloud

Model
Baseline-normalized term sentiment.
Data inputs
Around 500 English reviews per game, tokenised nightly. Terms are coloured by each term's positive share minus that game's own baseline positive share, so relative aspects surface rather than flattening against a fixed 50/50 threshold.
Validation
Not applicable. The cloud is descriptive, not predictive. It summarises what reviewers wrote, not what they will write.
Confidence
Higher for games with at least 200 English reviews. Lower for smaller review pools where individual reviewers can dominate the per-term positive share.

How we validate

Every published developer sales figure, every publisher wishlist disclosure is logged as a calibration anchor. When a new number lands and the model output falls outside a reasonable band, that becomes the seed of the next revision.

Calibration is continuous rather than periodic.

Where we make it visible

Where PGS labels a figure as an estimate, it is generated by one of the models described on this page rather than taken from a private or undisclosed source. Where a page shows a sales figure or a wishlist total, it is derived from the inputs listed on this page, not scraped from a private source.

If you have a disclosed sales or wishlist figure that differs materially from a PGS estimate, email Paul at paul [ at ] puredmg.com with the Steam AppID and the source. Verified real-world figures are valuable calibration data and help improve future model revisions.

What we do not model

Data freshness

Steam player counts are polled continuously. Retail prices are refreshed daily. Review analysis is refreshed nightly. Other source frequencies vary according to the availability and limits of the underlying service. Individual pages display their latest update time where applicable.

This page will grow as new metrics ship. If a number on the site is not described here, ask and it will be.