Scores

Why last year's 7 is this year's 7.5

A tool under commercial pressure has every reason to let its scale creep up and none to let it creep down.

By 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

Rating scales creep upward because every commercial incentive a tool has points the same way: low scores lose users, high scores get shared. The creep arrives through small, defensible changes rather than one decision, so last year's 7 can quietly become this year's 7.5. It is drift with a direction, and the direction is up.

What pushes it

A user who gets a low score is a user who leaves and does not come back to buy a second result. A user who gets a high score shares it, screenshots it, sends it to a friend who then tries the tool too. Neither of these facts requires anyone at the company to decide to inflate anything - they just mean that a tool tuned by watching engagement will trend toward whatever mapping keeps people scoring well, and that pressure only runs in one direction. The pattern is documented in human rating systems too: Filippas, Horton and Golden (2019) found marketplace ratings "prone to inflation", driven by raters' reluctance to harm the person rated, and argue that such systems "sow the seeds of their own irrelevance".

This is a different mechanism from ordinary calibration drift, where a rubric gets edited or a population shifts for reasons unconnected to outcome. Calibration drift can move a scale either way, because nothing about it is directional. Inflation specifically is drift with a business incentive attached to one direction, and it compounds in a way undirected drift does not: each small upward nudge makes the next nudge look smaller by comparison, and a tool can walk its whole population up a point over a couple of years without any single change looking like a decision.

How it compounds

The mechanism is gradual, not a single retuning event. A slightly generous rubric edit here, a rounding change that favours the higher of two borderline values there, a new axis added at a weight that happens to lift the total - none of these look like inflation in isolation, and each one is small enough to defend on its own merits. Stacked over enough small changes, the distribution the tool produces has moved without a single line in a change log describing it as movement.

This is also distinct from generosity bias, which is a static tilt a tool launches with and then holds. Generosity bias is a starting position. Inflation is what happens to that position, or to a strict one, once commercial pressure gets to keep nudging it for long enough.

How it shows up

Without inventing numbers, the shape of the effect is describable even where the size is not. A population under inflationary pressure produces fewer results in the low half of the scale over time, more results bunched near the top, and a shrinking gap between an average submission and an excellent one - because the tool has less room left to express the difference once most of its output already sits near the ceiling.

A reader without access to the tool's internal data cannot measure this precisely, but a few things are visible from outside: whether public examples of "typical" results have gotten visibly more favourable over a tool's lifetime, and whether long-time users report that a score which used to feel notable now feels ordinary. Neither is proof. Both are worth noticing.

What to do about it

There is no reader-side fix for inflation, because it lives inside the tool, not inside your submission. What you can do is treat any single score as tied to the version of the tool that produced it, not as a timeless value, which is the same discipline a recalibration requires for a different reason: a number from two years ago and a number from today are not automatically on the same scale, and inflation is one of the quieter ways that stops being true.

The honest tools are the ones that make this checkable at all. Rate Cock shows its axis scores on public entries rather than a single blended figure, which at least lets a reader watch whether the distribution across all six is shifting the same way a total alone would hide. Most tools give you no distribution to watch and ask you to trust that the number means the same thing it meant last year.

Inflation is a property of the business, not of your photo, and the geometry that moves a score for reasons unrelated to the tool's calibration is a separate and more immediate concern for any single result you get. The same caution applies outside automated tools too: a human reviewer working through a large volume of submissions is not immune to the same upward creep, for the same reason - a reputation for being harsh costs a reviewer business the same way it costs a product users. A property that cannot inflate this way is a purely physical one: a measurement in centimetres does not have a commercial incentive built into the number itself, which is one of the clearer differences between that category and a rating tool's.

Read next

Full archive