Methodology
PureGameStats publishes proprietary estimates for sales, wishlists, and pre-launch momentum. This page describes the inputs each model uses, how we validate its output, and where each estimate should be trusted versus taken as a directional signal.
Last methodology review: 3 September 2026
What we do not do: publish coefficients, weights, or the specific formulas that generate each number. Those are the parts of the model that took time to tune and remain proprietary. Everything else is documented below.
Sales estimate
- Model
- PGS Sales Model, current generation.
- Data inputs
- Steam review totals and positive-to-negative ratio, genre bands, tag dampeners for known outlier categories, launch cadence and days-since-release, regional review-language mix (used to detect regional-market titles that under-report reviews), Early Access graduation signal, viral-peak anchoring for mega-launches.
- Validation
- Continuously cross-referenced against developer-reported sales figures. Every new dev-published number is a calibration anchor.
- Confidence
- High for released commercial titles more than 30 days post-launch.
Medium for pre-release estimates and titles inside their launch window.
Lower for free-to-play viral hits where community gifting and platform-cross-play distort review counts.
Sales estimates should be treated as modelled ranges rather than audited sales figures. They are designed primarily for comparative analysis across titles and for estimating the likely order of magnitude of an individual game's performance.
Wishlist estimate
- Model
- A segmented model combining Steam wishlist rank and follower behaviour.
- Data inputs
- Steam public wishlist rank when disclosed by Valve, follower count, first-tag genre multiplier, Early Access flag, follower conversion rate for rank-stable titles, segment-straddle detection for fast movers.
- Validation
- Cross-checked against developer-disclosed wishlist totals whenever those become public.
- Confidence
- High for titles inside Steam's disclosed rank window.
Medium for long-tail ranks and titles with a stable rank but rapid follower growth.
Advisory for unranked titles with fewer than a few hundred followers.
Wishlist estimates for event-window debuts (Gamescom, SGF, Next Fest, TGA) tend to run below the eventual total because these games arrive via discovery outside a normal follower base. We do not retune the model against these events since it would drag accuracy in normal weeks.
Potential Hits momentum
- Model
- Composite score, ranked output.
- Data inputs
- Follower velocity, press-coverage volume (major gaming outlets), live Twitch presence, Steam public wishlist rank, Google Trends search interest, franchise-recognition, Discord community.
- Validation
- Retrospective, not predictive. We track how the score at T-30 days correlates with subsequent launch-week peak concurrent users. The board self-corrects as new launches age in and out of the window.
- Confidence
- Advisory ranking rather than a numeric forecast. Useful for spotting which unreleased and just-released games are showing signs the other lists do not surface.
Sudden viral hits remain a structural blind spot for any pre-launch model that reads Steam-native signals. PEAK, Schedule I and similar titles can go from very limited pre-launch signal to chart-topping CCU in weeks. We accept this and do not tune the model around exceptional viral events. However, the combination of inputs can still provides early indicators of breakout behaviour beyond the normal parameters.
Review word cloud
- Model
- Baseline-normalized term sentiment.
- Data inputs
- Around 500 English reviews per game, tokenised nightly. Terms are coloured by each term's positive share minus that game's own baseline positive share, so relative aspects surface rather than flattening against a fixed 50/50 threshold.
- Validation
- Not applicable. The cloud is descriptive, not predictive. It summarises what reviewers wrote, not what they will write.
- Confidence
- Higher for games with at least 200 English reviews. Lower for smaller review pools where individual reviewers can dominate the per-term positive share.
How we validate
Every published developer sales figure, every publisher wishlist disclosure is logged as a calibration anchor. When a new number lands and the model output falls outside a reasonable band, that becomes the seed of the next revision.
Calibration is continuous rather than periodic.
Where we make it visible
Where PGS labels a figure as an estimate, it is generated by one of the models described on this page rather than taken from a private or undisclosed source. Where a page shows a sales figure or a wishlist total, it is derived from the inputs listed on this page, not scraped from a private source.
If you have a disclosed sales or wishlist figure that differs materially from a PGS estimate, email Paul at paul [ at ] puredmg.com with the Steam AppID and the source. Verified real-world figures are valuable calibration data and help improve future model revisions.
What we do not model
- Twitch viewer counts, streamer counts, and hours watched are direct pass-throughs of Twitch's own live data. No modelling.
- Steam player counts are direct polls of Valve's public numbers.
- Prices are the current retailer price as scraped from that retailer's product page.
- Discord member counts are direct queries of Discord's invite endpoint.
Data freshness
Steam player counts are polled continuously. Retail prices are refreshed daily. Review analysis is refreshed nightly. Other source frequencies vary according to the availability and limits of the underlying service. Individual pages display their latest update time where applicable.