Scores
How an internal value becomes a number out of ten
Between whatever the system computes and the number you see there is a mapping step, and it decides more about your score than the computation does.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
Normalisation is the step that turns whatever a rating tool computes internally - a probability, a distance, a raw model output - into the number out of ten you see. The tool never computes "7.4" directly. That mapping step is where most of what makes two tools disagree actually happens.
Why the step exists at all
Whatever a model computes internally rarely arrives in a convenient range. It might be a probability between 0 and 1, a distance from some reference point, or an unbounded score with no natural ceiling. None of those are "out of ten" on their own, so a tool has to decide how to convert one into the other. That decision is normalisation, and every tool that shows you a number out of ten has made it, whether or not it tells you how.
The computation that produces the internal value in the first place is a separate stage, sitting one layer below this - what a vision model is actually doing with a photo is a different subject, and it stays roughly the same across most tools in this category. Normalisation is what happens after that shared layer, in the part each service actually built.
The choices inside the mapping
Linear. The internal value is stretched or compressed to fit the target range directly - a straightforward multiply-and-shift. Simple, and it inherits whatever shape the internal value already had, crowding wherever the underlying distribution crowds.
Clipped. Values above or below some threshold get pinned to the ends of the scale rather than extrapolated. This is why very few tools ever show a 1 or a 10: the clip point is usually set well inside the theoretical extremes, because a tool that regularly prints its own floor or ceiling looks unstable.
Percentile-based. The internal value is compared against the tool's own history of past results, and the score reflects where this submission falls in that distribution rather than any fixed rubric maximum. This is a specific and different kind of mapping worth its own treatment - a percentile score changes meaning as the population changes, even if your submission does not.
Each of these produces a number that looks identical on the page. Nothing about a 7.4 tells you which of the three produced it.
Why this is where calibration lives
Calibration is often described vaguely as "how accurate a tool is," but at the mechanical level it is specifically this mapping step. A tool is well calibrated when its normalisation produces numbers that mean the same thing across submissions and across time - a 7 today implies roughly what a 7 meant last month, on the same tool. A tool can compute a perfectly reasonable internal value and still be poorly calibrated, if the mapping onto the visible scale is inconsistent, undocumented, or silently retuned. The machine-learning version of this step is well studied: Guo et al. (2017) found that temperature scaling, a post-processing mapping with a single parameter, was "surprisingly effective" at calibrating modern networks on most datasets - a sign that calibration is often repaired after the model rather than inside it.
This is also why two tools computing genuinely similar internal judgements can print very different numbers for the same submission. The disagreement is rarely in the underlying model. It is almost always downstream of it, in a mapping step neither tool discloses. The fuller account of why two tools disagree on one photo traces that gap back to exactly this stage.
What a reader can actually check
Almost nobody publishes their normalisation function. What some tools do publish, or make inferable, is distribution information - a public board of results, a histogram, anything that shows where a given number sits relative to others. That is not the mapping function itself, but it is the next best thing: it lets you see the shape normalisation produced, even without seeing the rule that produced it.
A tool that shows nothing at all - no distribution, no rubric, no stated scale - is asking you to trust a mapping you cannot even indirectly observe. Rate Cock publishes per-axis scores on public entries, which at minimum lets you see the shape each axis lands in across many submissions, rather than trusting a single blended figure with an invisible mapping behind it.
What normalisation is not
It is not the computation itself, which sits below it and is shared across most of the category. It is not a measurement in any physical sense - a length in centimetres comes from a tape and a method, not from a mapped model output, and normalisation has nothing to say about that kind of number. And it is not what a human reviewer does with an impression, which does not pass through anything resembling a numeric mapping step at all.
Normalisation is specifically the scale-layer decision that turns whatever a rating tool computed into the digit you read. It is invisible on the page and it is doing more work than the model underneath it, which is exactly why it is worth knowing it exists even when you cannot see inside it.