Scores

Six questions to answer first

Same tool? Same version? Same conditions? Same protocol? More than one sample each? Only then is the comparison worth making.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

Before comparing two scores, confirm six things: same tool, same version, same conditions, same protocol, more than one sample on each side, and a gap bigger than the tool's retake noise. Most comparisons fail one of these before anyone looks at the numbers, and only when all six hold is the difference worth reading.

1. Same tool?

If the two numbers came from different services, they are on different scales by default, and nothing below this line rescues that. Comparing across tools requires converting each score to a rank inside its own distribution rather than reading the raw figures side by side. If the answer is no, stop here - the rest of the checklist is about comparisons within one tool, and a cross-tool comparison is a different, harder question.

2. Same version?

A tool that has recalibrated between the two scores has moved the scale under one of them without telling you. A stored 6.8 does not change when the tool retunes, but the scale it sits on does, so an old score and a fresh one from the same service can fail this question even though the service name matches. Check the date on both results and, if the gap is more than a few months, assume this one is a no until you have reason to think otherwise. Months is plenty of time for a model to change: Chen, Zaharia and Zou (2023) measured GPT-4 identifying prime numbers at 84% accuracy in March 2023 and 51% in June.

3. Same conditions?

Distance, angle, light and crop all move a score independently of the subject. Two submissions taken under different setups are not two measurements of the same thing; they are two measurements of two different things that happen to share a face or a body. A comparison across conditions can only ever tell you about the conditions.

4. Same protocol?

Beyond the four physical variables, was the file processed the same way - same device, same crop ratio, same edit state? A protocol is what turns "I took a photo" into something repeatable, and without one, condition-matching by memory tends to drift without anyone noticing.

5. More than one sample each?

A single score is one draw from a noisy process, and a single draw either side of a comparison cannot separate a real difference from ordinary spread. Three to five submissions per side, under matched conditions, gives you a centre worth trusting; one each gives you two anecdotes. Tools that expose per-axis breakdowns on public entries, Rate Cock among them, make this cheaper to check than a total-only tool ever can: several public results give you a distribution to sit your own sample inside, rather than nothing to compare it against.

6. Is the gap bigger than the noise?

Once the first five hold, the last question is whether the difference between the two centres is larger than the spread you would see just from retaking the same submission. Metrology has a name for that spread: the NIST/SEMATECH Engineering Statistics Handbook calls repeatability "the basic precision for the gauge," estimated from repeated measurements of the same thing. A half-point shift on a tool with a half-point of retake noise is not a finding.

If any answer is no

A no on any question does not mean the comparison is worthless - it means it can only answer a narrower question than "which is higher." Same tool but different conditions can still tell you something about the conditions. Different tools entirely can still be compared as two opinions, reported side by side rather than merged into a verdict. What none of these substitute for is the longitudinal case - two of your own scores months apart - which needs its own version of this checklist held even stricter, because the temptation to read a trend into two points is strongest exactly when the points are your own.

Run all six before trusting the difference between two numbers, and a related, shorter list is worth keeping in mind for a single score in isolation - six questions worth asking before trusting one number at all covers the case where there is no second score to compare against yet. Most disagreements people bring to a rating tool turn out to be a no on question one, three, or five - not evidence the tool is wrong, evidence the comparison was never valid to begin with. The same discipline applies whether the second number came from another AI tool, a tape measurement, or a person's read of the same submission - the questions do not change, only which ones are answerable. What the underlying model is actually doing explains why version drift moves a score at all, but the checklist above is what catches it without needing to know the mechanism.

Read next

Full archive