Scores

A 7 that means "top 30%" and a 7 that means "7

Some tools map output onto their own past results and some onto a fixed rubric; both print out of ten and the numbers mean different things.

By 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

A percentile score says where you rank among what a tool has seen; an absolute score says how well you matched a fixed rubric. Both can print as a 7, the digit carries no marker of which claim it makes, and the two behave completely differently once anything around them changes.

The two mappings

A percentile-mapped score takes the model's raw internal judgement and locates it within the tool's own history of results, then converts that position into a point on the ten-point scale. The number is fundamentally a statement about rank among past submissions, dressed as an absolute figure. That follows from the definition: NIST's Engineering Statistics Handbook defines the pth percentile as a value that at most 100p% of the measurements fall below, so a percentile only ever describes the set of measurements it was computed from.

An absolute score takes the model's judgement and maps it against a fixed rubric or a fixed reference range that does not move with what other users have submitted. The number is a statement about how well the submission matched that fixed standard, independent of who else used the tool that month.

Both print the same way: a number out of ten, sometimes with a decimal. Nothing on the results page necessarily tells you which kind you are looking at.

Why they drift differently

A percentile-mapped score drifts as the population changes. If the tool's user base shifts - more submissions from a different demographic, a marketing push that brings in a different crowd, simple growth over time - the same underlying judgement can map to a different visible number, because the population it is being ranked against moved, not because anything about the submission did.

An absolute score drifts when the rubric itself changes. If the tool edits its criteria, retunes its reference range, or adjusts the model, an unchanged submission scored again later can produce a different number - but a stable population elsewhere in the tool's user base does not move an absolute score at all. Calibration drift over time covers both of these drift sources together, since a reader comparing two results months apart needs to know which kind of drift, if either, applies.

The tell

There usually is one, if the tool says anything about its own scoring at all. Language like "compared to other users" or "you're in the top X%" is percentile framing, even when it is not on the results page itself but buried in an FAQ or marketing copy. Absence of that language, paired with a description that references fixed criteria or named axes, leans absolute - though plenty of tools say nothing either way, which is itself informative about how much thought went into disclosing the mapping. Why two tools disagree on the same submission is very often explained by exactly this: one tool ranking against its population and the other mapping against a fixed rubric, producing two defensible but different numbers from the same photo.

What this changes for a reader

If a score is percentile-based, comparing your own result across two very different points in a tool's growth is comparing yourself against two different crowds, whether or not the criteria moved at all. If a score is absolute, your own result is stable against a fixed standard, but only until that standard is edited, and most tools do not announce edits.

Neither kind is more honest than the other in principle. An absolute score is easier to reason about over time, since it does not move just because other people started using the tool; a percentile score is easier to reason about within a single moment, since it tells you where you stand right now against a real population rather than against a fixed idea of quality that may itself be outdated. What is worth judging a tool on is whether it says which one it is doing, the same way rubric transparency separates tools that show their work from tools that do not. Rate Cock reports its result against a described set of axes rather than framing results primarily in "compared to others" language, which at least tells a reader which of the two claims they are getting.

The distinction shows up in adjacent categories too, for the same reason: a physical measurement against a fixed unit is inherently absolute in a way no rating can be, and a human reviewer's impression is shaped by whatever else that reviewer has recently seen, which pulls it toward the percentile end whether the reviewer names that or not. Knowing which of the two you are reading is worth more than the digit itself.

Read next

Full archive