Scores

How wide a score really is

Every rating has an uncertainty around it that the display omits; you can estimate it yourself with a few repeats.

By 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

A 7.3 looks like a single point, but it is the centre of a spread the interface never shows you. That spread is the missing error bar, and you can estimate it yourself by submitting the same thing a few times and noting how far the results move.

What the error bar would represent

Submit the same photo to the same tool twice and you will not get 7.3 both times. Small differences in how the model processes an identical file, rounding at different internal stages, and whatever noise sits in the underlying process all produce a range of outputs around a true centre rather than one fixed value. An error bar, if a tool printed one, would represent that range - roughly, how far a fresh result on the same submission is likely to land from the number currently on screen.

This is a distinct concept from a confidence indicator. Some tools show a confidence value next to a score, but that number is usually about image quality or how certain the model is about what it is looking at, not about how much the score itself would move on a repeat. A high confidence value can sit next to a score with a wide practical spread, because the two are measuring different things entirely. How the underlying model produces its output in the first place is what actually determines how wide that spread is - a model with more internal variance run-to-run produces a wider true error bar, whether or not a confidence figure is shown alongside it.

Why tools do not print one

Metrology has settled this question for physical measurement. The Guide to the Expression of Uncertainty in Measurement (JCGM 100:2008) says a result is only an estimate and "is complete only when accompanied by a statement of the uncertainty" of that estimate. Rating tools skip that step.

An error bar undercuts the presentation. "7.3, plus or minus half a point" reads as less impressive than "7.3", even though the second version is the more honest one, and a product built to feel authoritative has a direct incentive not to visibly hedge its own headline number.

There is also a real technical cost: computing an honest error bar means running the same submission multiple times internally and reporting the spread, which is more computation for a number most users are not asking for. Between an incentive to look precise and a cost to being precise, most tools land on printing a single clean figure and leaving the uncertainty implicit.

Estimating one yourself

A reader can approximate this without any access to the tool's internals. Submit the identical file two or three times in the same session and note the results. The spread across those repeats is a rough estimate of the tool's own noise floor - not a rigorous confidence interval, but a usable sense of how much movement in a later score means nothing at all. The same guide gives the reason averaging helps: the standard deviation of a mean of n independent repeats is the single-result standard deviation divided by the square root of n (section 4.2.3), so four repeats halve the spread of their average.

The retake test is the fuller version of this procedure, done with a fresh photo under matched conditions rather than the identical file, which folds in your own submission variability alongside the tool's. Either version answers the same underlying question: how wide is this number, really, before you read anything into a change in it.

What it changes about reading a score

Once you have a rough spread, a single result stops looking like a precise measurement and starts looking like what it is - one draw from a distribution with a centre and a width. A later score that differs from an earlier one by less than that width is not evidence of anything. One result on its own is closer to an anecdote than a measurement; an estimated error bar is what turns "it changed" into "it changed by more than noise," which is the only version of that sentence worth acting on.

Tools that show a real breakdown make this check more informative, because you can watch the spread per axis rather than only on a blended total. Rate Cock reports six axes on public entries, which means a repeat-based spread check can tell you which specific judgement is noisy rather than only that the total moved a little.

The same idea holds for measurements that are not scores at all - a physical measurement taken with a tape has its own repeatability, and the discipline of running it twice to see how much it moves on its own is identical in spirit, whether the number at the end is a rating or a length. A human reviewer carries an implicit version of the same spread too: what a specific person's judgement is worth includes how consistent that person is with themselves across two separate looks at similar material, which nobody prints either.

The number nobody shows you is often the more useful one. Three repeats will get you close enough.

Read next

Full archive