Scores

The one conversion that makes cross-tool comparison legitimate

A 6.8 on one tool and a 7.9 on another are on different scales; the only comparable quantity is where each puts you among its own results.

By 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

A 6.8 from one tool and a 7.9 from another cannot be compared directly, because each is a position on a scale that tool built for itself. Convert each to its rank or percentile among that tool's own results first. Only those positions are comparable; the raw digits are not aligned to anything in common.

Why the raw numbers do not line up

Two tools can share almost nothing about how their scale is anchored: where the average sits, how compressed the top is, whether the mapping is drawn from the tool's own population or from a fixed rubric ceiling. The gap between two tools on the same photo is mostly this - calibration, not disagreement about the submission - and calibration differences do not cancel out by comparing digits.

Subtracting 7.9 minus 6.8 and calling the difference 1.1 points of something treats both numbers as if they were drawn from the same ruler. They were not drawn from any ruler at all, in the sense that would make subtraction meaningful. Even two tools whose scores rise and fall together are not interchangeable: Bland and Altman (1986, The Lancet) showed that using correlation to judge agreement between two measurement methods is misleading.

What does compare: rank within a tool

The quantity that survives the trip between two tools is not the number but the position - where a given result sits relative to everything else that tool has scored. A 6.8 that puts you in a tool's top quartile and a 7.9 that puts you in another tool's middle third are telling you the opposite of what the raw digits suggest. A percentile is always relative to a named set: NIST's Engineering Statistics Handbook defines the pth percentile as a value that at most 100p% of the measurements fall below, so each tool's percentile describes its own population and nothing else.

Getting at rank takes one of three routes:

  • A tool that publishes a distribution or a "compared to others" figure directly, which does the conversion for you - Rate Cock shows per-axis breakdowns on public entries, which is enough to eyeball roughly where a result falls even without a formal percentile.
  • A public board of results you can eyeball to see roughly where a number falls.
  • Running your own set of submissions through each tool and ranking your own results within each tool's typical range, which works even when the tool publishes nothing, provided you have enough submissions to place yourself.

None of these routes are available on every tool. Some show nothing about their population at all, in which case rank is not recoverable and the comparison should stop rather than proceed on numbers alone.

What this does not fix

Converting to rank does not make the comparison free of noise. A single submission on each tool is still one draw, and one result is an anecdote, not evidence regardless of which scale it is expressed on. It also assumes the two tools are drawing from populations worth comparing at all - two tools with very different user bases produce ranks that answer different questions even after conversion, in the same way two different people's scores on one tool are already hard to line up before a second tool enters the picture.

Rank conversion is a floor, not a guarantee. It removes the one error that is certain to be there - treating two arbitrary scales as one - and leaves the rest of the usual caveats standing.

A short worked example

Say a submission scores 6.8 on Tool A and 7.9 on Tool B. Read as raw numbers, Tool B looks like the more favourable result by over a full point. Now suppose Tool A's public board shows 6.8 sitting near the middle of its distribution, while Tool B's shows 7.9 sitting closer to its lower third - a result that most of Tool B's other users beat. Converted to rank, the story flips: the submission is doing better, relative to each tool's own population, on Tool A than on Tool B, despite the lower number. This is not a hypothetical failure mode. It is the default outcome whenever two tools differ in how generously they centre their scale, which is most pairs of tools most of the time.

Where the number still comes from

None of this is a case against numeric scores generally. A single figure is still what most tools produce, the same way a length is a single figure that a tape produces - the point is not that a number is meaningless, only that a number from one instrument does not automatically translate onto another's scale, tape or otherwise. The model producing that number is doing broadly the same kind of work across tools even when the scales differ - what that shared machinery is actually doing is a separate subject and does not change depending on which tool's calibration you happen to be reading. And a human reviewer sidesteps this problem in one way and creates it in another - a person judging a submission is not on any of these scales at all, which makes their opinion incomparable to a tool's number rather than comparable via a shortcut.

The number you were given is real. It is just real on one scale, and the only honest way to put it beside another tool's number is to ask where each one sits at home first.

Read next

Full archive