Scores

The mean of two incompatible numbers

Averaging a 6.5 and an 8 from two tools produces a number on no scale at all.

By Updated 3 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

Two tools, two scores, and the instinct to split the difference. A 6.5 and an 8 become a 7.25, and the 7.25 feels like a more balanced verdict than either input alone. It is not a verdict at all - it is arithmetic performed on two numbers that were never measuring the same thing.

Why the mean does not work here

Averaging assumes both numbers sit on a shared scale, with the same zero, the same ceiling, and the same distribution of results underneath them. Reliability research separates exactly these two cases: Koo and Li (2016) distinguish raters whose scores are merely correlated "in an additive manner" - one consistently higher than the other - from raters who give identical scores, and only the second kind supports treating their numbers as interchangeable. Two tools rarely share any of that: one may centre its results around six, another around seven and a half; one may compress its top end, another may not. A 6.5 from a strict, generously-spread tool and an 8 from a lenient, ceiling-compressed one can represent almost the same underlying impression, or wildly different ones - there is no way to tell from the numbers themselves, so there is no way to tell what the average of them represents either.

Averaging temperatures reported in Celsius and Fahrenheit produces the same failure for the same reason: the numbers look commensurable and are not, and the mean of two incommensurable numbers is not a smaller error, it is a new number with no referent. That the two tools may run broadly similar vision models underneath does not rescue the average - the model is one layer, and the rubric and scale built on top of it are what actually differ between services.

What the average hides

A blended figure also erases the disagreement itself, which is usually the more useful signal. If two tools land far apart, that gap is telling you something - about calibration, about a rubric difference, about which one is more generous - and averaging it away trades that information for a single tidy digit that answers nothing.

What to do instead

Convert each score to a rank or a percentile within its own tool's distribution first, if the tool publishes one, and compare the two positions rather than the two raw numbers. Where no distribution is published, the two numbers can only be reported side by side, with the tool named against each one. Comparing scores across tools by converting to rank is the whole procedure, and it works precisely because it never lets the two raw scales touch each other directly.

Rate Cock's public entries expose the breakdown behind each total, which is a start toward this: a rank derived from six visible axes is comparable in a way a bare average of two black-box totals never is. The same caution applies off a screen entirely - averaging a length in centimetres with an AI aesthetic score is the same error again, just with a tape measure standing in for the second tool, and a human judge's read is a third scale that cannot be folded into either.

The fix is not a better formula. It is refusing to average two things that were never denominated the same way, and reporting them as what they are: two separate opinions from two separate scales.

Read next

Full archive