Scores

A single number cannot carry the weight put on it

One rating is a draw from a distribution; the useful quantity is the centre of several, and the useful question is how many is several.

By Updated 3 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

One result is not evidence because it is a single draw from a distribution, and it cannot separate signal from noise. Three to five repeats under the same conditions give a centre worth reading. A lone number only feels final because it arrives with nothing beside it.

Why a single draw is not enough

Every rating result carries some amount of noise - from the model, from small variations in how a photo was processed, from whatever the tool's internal generation happens to do on that particular run. A single number cannot separate the signal from that noise, because a single number is the sum of both with no way to tell how much of each is in it.

This is not a criticism of any specific tool. It is a property of any process that produces a value with variance, rating tool or otherwise: how a model actually generates that value is a separate question from how many times you need to sample its output, but the two are related - a noisier underlying process is exactly the kind that needs more repeats before its centre settles. One sample tells you where that sample landed, not where the underlying value tends to sit. The centre of several samples is a much better estimate of that underlying value than any one of them alone, for the same reason a single coin flip tells you nothing about whether a coin is fair and twenty flips start to.

Why three to five is a reasonable minimum

There is a point of diminishing returns, and it arrives sooner than people expect. Going from one result to three does the most work: it converts a single unverified draw into a rough centre and gives you a first sense of the spread around it. Going from three to five tightens that centre further. The arithmetic behind that is standard. The Guide to the Expression of Uncertainty in Measurement (JCGM 100:2008) gives the standard deviation of a mean of n independent observations as the single-observation figure divided by the square root of n, so four repeats halve the uncertainty and it takes sixteen to halve it again. Beyond five or so, each additional repeat buys progressively less, because the estimate has already mostly settled and the remaining uncertainty is not going away no matter how many more times you run it.

This is a directional claim, not a precise formula - the exact number that is "enough" depends on how noisy a given tool is, which varies and is not something a reader can know in advance. Three to five is a reasonable working minimum before treating a centre as meaningful, not a guarantee that five repeats have eliminated the uncertainty.

What stays unknown even then

Repeats settle the centre of the tool's output for one submission under one set of conditions. They do not tell you whether that centre is accurate in any absolute sense, whether the tool's calibration matches any other tool's, or whether the population it is implicitly scoring you against is one worth being scored against. A single result on its own is the weakest version of this problem; more repeats fix the sampling issue specifically and leave every other question about the score exactly where it was.

Repeats also do not fix a bad submission. If the underlying photo does not follow a standardised protocol, five repeats of an unstandardised photo just give you a well-estimated centre of a result that was never going to be comparable to anything else in the first place.

Reading the resulting centre

Once you have a centre from several repeats, that is the number worth remembering, not the first result you happened to get and not the highest one out of the batch. Rate Cock keeps a visible history per account, which makes running and reviewing a small series easier than it is on a tool that only shows the most recent result and discards the rest.

The same logic applies wherever a single judgement gets treated as a verdict. A physical measurement benefits from more than one careful attempt for exactly the same reason a rating does - hand position and tension vary between attempts the way model noise varies between runs. And a single opinion from one human reviewer is a sample of one in a different sense: it is one person's taste on one day, not a settled verdict, however confidently it gets delivered.

One number is a starting point. A settled centre from several is closer to something you can act on.

Read next

Full archive