Game review scores use a wide numeric range but occupy a small part of it in practice. The compression is structural, and understanding it makes scores more useful rather than less.

The sample is filtered before anyone plays

Outlets review a small fraction of what is released, and they choose titles with an audience, which means funded projects that reached completion and shipped through a publisher.

Games that fail earliest, during funding or production, never reach a reviewer at all, so the worst possible outcomes are largely absent from the distribution.

What remains is a set of competently made products differing in ambition and execution, and competent products legitimately cluster.

Scales are compared against other scores, not against a definition

Numbers acquire meaning through use. Once an audience reads a middling score as a warning, it functions as one regardless of what the scale was intended to express.

Reviewers write knowing that, so a genuinely average game receives a score that signals average within the shared convention rather than at the arithmetic midpoint.

The scale has effectively been redefined by its readers, and any single outlet that resists is simply misread.

Aggregation amplifies the effect

Score aggregators average across outlets, and averaging pulls results towards the centre of whatever range reviewers actually use.

Because those averages are cited in marketing and sometimes tied to contractual bonuses, the stakes attached to individual numbers rise considerably.

Higher stakes make outlets more careful about outliers, which tightens the band further, since an unusual score now requires far more justification than a conventional one.

Timing shapes what gets scored

Reviews are written near launch, when a game is at its least finished and its audience most eager, and the schedule leaves little room for extended assessment.

Games that improve substantially over the following year rarely get rescored, and games that decay through neglect keep the number they earned in the first week.

The score therefore describes a moment rather than a product, which is a limitation of the format rather than a failure of the reviewer.

The text carries the information the number cannot

A score compresses a long assessment into one value, discarding exactly the detail a reader needs to know whether the game suits them.

Two games with the same number can differ completely in what they do well, and the review body is where that distinction lives.

Readers who use scores as a filter and the text as the decision get most of the value, while readers who compare numbers across outlets are measuring conventions rather than games.