Scores

What can and cannot be compared, and how

Across tools, across time, across people, across photos, the complete map of which comparisons hold and what each needs to hold.

8 min readScores

A rating tool hands back a single number, and the number invites a comparison almost by reflex - higher or lower than what, than whom, than last time. Most of those comparisons are invalid the moment they are made, not because the numbers are wrong but because the two numbers were never eligible to sit next to each other. This is the map: four kinds of comparison, what each one requires to hold, and where each one breaks - and the approach this particular site takes to comparing tools against each other follows the same logic throughout.

Comparison one: across tools

Two different services, one submission, two different numbers. Tools disagree on the same photo mostly because of calibration, not accuracy - one tool maps its output onto its own historical distribution, another maps from a fixed rubric, and both print a figure out of ten that looks like the same kind of thing.

What this comparison needs to hold: nothing about the raw numbers, because the raw numbers are not the comparable quantity. Converting each score to a rank or percentile inside its own tool's distribution is the only version of this comparison that works, and it only works if the tool publishes enough results to build a distribution from. Where it does not, the honest move is to report both numbers with their tools named and stop there - not to average them, because averaging two different scales produces a number on no scale at all.

Where it fails hardest: when the two tools share a backbone. Many services sit on similar hosted vision models, so two tools agreeing is sometimes one judgement counted twice rather than independent confirmation, and two tools disagreeing despite a shared backbone is a stronger signal than the same disagreement between two unrelated tools.

Comparison two: across time

The same tool, the same person, months apart. This is the comparison most people actually want to make - "am I scoring higher than I used to" - and it needs three things to have held constant that usually have not: the tool's calibration, the submission conditions, and the subject's own state.

A tool that has recalibrated between the two dates has moved the scale under one of the scores without saying so, and rubrics do get edited, mappings retuned, and populations shift over a year without any announcement. Conditions drift quietly even when nobody meant them to - a little closer, a slightly different lamp, a looser crop - and each of those moves the score independently of anything the comparison is meant to be about.

What this comparison needs to hold: same tool, same version if checkable, same protocol, and ideally a control shot - one fixed reference image resubmitted alongside each real one, so that if the control moves, you know the tool moved and can subtract that from the real result. Without a control, a shift of a few tenths a year apart is unreadable: it could be the subject, the conditions, or the tool, and there is no way from the numbers alone to say which.

Comparison three: across people

Two different people, ostensibly the same tool, ostensibly comparable numbers. This is the weakest of the four kinds, and it is weak for a reason that has nothing to do with either person: two submissions differ in conditions, device, and protocol before either subject is even considered, and those differences are usually larger than any real difference between the two people.

Unless both people submitted under a genuinely matched protocol - same device family, same distance, same angle, same light, ideally on the same day - a comparison between their two numbers is mostly a comparison between their two photography setups. Even matched conditions leave the deeper problem: a rating tool orders submissions within its own training population, and "who scores higher" says less about either person than it says about where each one happens to fall in that population's spread.

What this comparison needs to hold: matched conditions at minimum, and an acknowledgment that even a well-matched comparison is answering "who did the tool prefer today," not a durable fact about either person. This is also the comparison most likely to be requested casually and answered with false confidence, because the two numbers exist and the temptation to read them at face value is strong.

Comparison four: within a set

Multiple submissions from one person, one session, ranked against each other. This is the comparison a rating tool is structurally best at, because everything that makes the other three fragile - device, calibration, population - is held constant by construction: it is the same tool, the same day, the same person, usually the same rough conditions.

A tool orders a set of your own submissions more reliably than it scores any single one of them, which sounds like a small distinction and is not: the absolute number carries all the calibration uncertainty described above, while the relative order within one session mostly cancels it out, because whatever the tool's scale is doing, it is doing the same thing to every photo in the set.

Where this comparison still fails: at the edges of the set, where near-ties are noise rather than signal, and when the images in the set were not actually taken under matched conditions - a mixed-condition set has no single condition to attribute its order to, and its internal ranking is then just as uninterpretable as any of the comparisons above.

What each comparison actually requires, side by side

Comparison Requires Fails when
Across tools Convert to rank/percentile, never raw numbers Tools share a backbone, or one has no published distribution
Across time Same tool version, same protocol, a control shot Rubric edited, mapping retuned, conditions drifted
Across people Matched device, distance, angle, light Any condition differs, or population context is ignored
Within a set Same session, roughly matched conditions Set is mixed-condition, or near-ties are read as real gaps

Reading this table top to bottom is also reading the comparisons in order of how much has to hold for each one to work - across tools needs the least in common between the two numbers and is correspondingly the weakest comparison to make casually; within a set needs the most in common by construction and is correspondingly the strongest.

A worked example of a comparison that looks valid and is not

Take two results: a 7.1 from six months ago and a 7.6 today, same tool, same account. On the surface this passes the same-tool test and looks like a half-point improvement worth noting. Two things could still be hiding inside that gap. First, the tool may have recalibrated in the interim - a routine model or mapping update that neither user nor result page ever announced - in which case the scale under the 7.6 is not the scale under the 7.1, and the half point is partly or entirely an artefact of that shift. Second, the two submissions may not share a protocol: a slightly closer camera, a different light source, a looser crop. Either factor alone can produce a half-point move with nothing about the underlying subject having changed at all. The only way to tell the difference between a real change and either artefact is the control shot described above - a single fixed reference resubmitted alongside both real attempts, whose own movement (or lack of it) tells you how much of the half point belongs to the tool and the conditions rather than to anything else.

The common thread

All four comparisons fail for versions of the same reason: a rating tool's number is not portable outside the conditions it was produced under, and every comparison either controls for those conditions or inherits their noise. Cross-tool comparisons fail on calibration. Cross-time comparisons fail on drift. Cross-person comparisons fail on device and population. Within-set comparisons are the one case where the conditions are mostly shared already, which is why they are the one comparison worth trusting without a checklist.

A short checklist exists for the general case - same tool, same version, same conditions, same protocol, more than one sample each - and running it before trusting any two-number comparison catches most of the failures above without needing to diagnose which specific one applies.

What sits outside this map entirely

Two comparisons are worth naming as categorically different rather than merely harder. A comparison against a physical measurement is not on this map at all - a tape and a rating tool are answering different questions, and no amount of matched conditions makes a centimetre figure and an aesthetic score comparable. A comparison against a human reviewer's read is similarly a different kind of output, a response in a register rather than a position on a scale, and treating it as one more data point to average against a tool's number repeats the averaging error from a different direction.

The map, compressed

Across tools: convert to rank, never average, distrust agreement between tools sharing a backbone. Across time: match the tool version and the protocol, use a control shot, expect drift you have to account for. Across people: match conditions at minimum, and read the result as "who the tool preferred today," not a durable ranking. Within a set: the one comparison the tool is built for, provided the set itself was taken under one protocol.

None of this requires knowing what the underlying model is doing - the failures above sit entirely in the layer above the model, in rubric, calibration, and protocol, which is also where every fix lives. What holds across all four is simpler than any of the individual rules: a number only compares to another number that shares its conditions, and the work of comparing two scores is mostly the work of checking whether that was ever true. Where a tool shows its work - a breakdown, a version, a distribution of public results, the way Rate Cock does on its entries - checking gets cheaper. Where it does not, the comparison has to be made more cautiously, or not made at all.

Read next

Full archive