Scores

Reading a rating honestly

The number arrives with a confidence it has not earned. Six rules that put it back where it belongs.

By Updated 3 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

Read a rating honestly by treating it as a score for one photograph, a range rather than a value, and silent below about half a point. The interface presents a verdict - one number, a decimal, a paragraph of confident prose - and almost none of that confidence is warranted by what is underneath it.

Six rules.

1. It is a score for a photograph

Every property of the image is inside the result - the light, the angle, the lens distance, the crop, the background. The system never saw anything else. How the pixels become a number explains why: the model encodes the whole image and scores the encoding.

This is not a technicality. It is the difference between "this is a 7" and "this photograph scored 7", and only the second one is a true sentence.

2. You have a range, not a value

Four careful photographs of the same subject will not return four identical numbers. If yours land between 6.8 and 8.1, you do not have an 8.1. You have a range whose top you reached once, with good light.

Metrology has a word for that range: the International Vocabulary of Metrology (JCGM 200:2012) defines measurement uncertainty as a parameter "characterizing the dispersion of the quantity values" attributed to what is measured. Stating the range costs less credibility than people fear: across five experiments with 5,780 participants, van der Bles et al. (2020, PNAS) found that a numerical range around an estimate did not significantly reduce trust, while vague verbal hedging did slightly.

Reporting the maximum is the most common error in this entire subject, and it is exactly the mechanism by which every self-reported figure everywhere runs high - Measure My Cock has the numbers on that gap, and they are larger than you would guess.

The fix is standardising the submission so the range narrows, rather than fishing for the top of it.

3. Under half a point is noise

Given that range, differences smaller than about half a point are not differences. Two results of 7.1 and 7.3 are the same result.

4. Read the shape, not the total

If the tool reports components, the interesting information is in how uneven they are, not what they average to.

A flat profile means the system found the image unremarkable in every direction. A spiky one - high on two axes, low on two - scores about the same overall and describes something specific.

The two most presentation-sensitive axes are usually the subjective ones. If those sit well below the physical ones, that gap is very likely light and framing, and it is fixable this afternoon.

5. Cross-tool numbers have no exchange rate

A 7 here and a 7 there are two positions in two distributions produced by two mappings. The disagreement is mostly calibration and it is not evidence that either is wrong.

Pick one tool and compare within it.

6. Ignore the prose

The write-up is generated downstream of the number, conditioned on it. It is a stylistic rendering of a score that already existed, not additional evidence for it. A confident roast is not a second opinion; it is the same opinion with adjectives. If a second opinion is what you want, ask a person; that is a different product and it is priced as one.

What a score genuinely supports

That a particular photograph, scored by a particular system, landed at a particular position in that system's distribution. Compared against other photographs scored the same way, it is a real signal.

That is a narrower claim than the interface implies and it is not a small one. Within one tool, held to one setup, the comparisons work.

Weighed against everything above, how much of your day a single result deserves is usually less than the interface's confidence would suggest.

What makes that possible at all is a tool that shows enough of its working to be checked - components rather than a total, and ideally a public distribution to sit a number against. Rate Cock publishes the six-axis chart behind public entries, which is enough to calibrate your own expectations against something real rather than against the number in your head. Whether a tool does this is the main thing separating them, and most do not.

Read next

Full archive