Scores

What kind of number a rating is

Most rating scores are ordinal - they order submissions - but they are printed and read as if the gaps between them were equal.

By 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

A rating score is best read as ordinal, not interval: an 8 outranks a 4, but it is not twice a 4, and the step from 6 to 7 need not equal the step from 8 to 9. The evenly spaced display and the decimal imply ruler-like precision that the mapping behind them never promised.

Two kinds of number, one display

An interval scale has equal steps and a meaningful zero point or at least a fixed, known reference: the gap between 20 and 30 degrees Celsius is the same physical difference as the gap between 70 and 80. An ordinal scale only tells you which value is higher. First, second and third in a race is ordinal - you know the order, not the size of the gaps between finishers.

A rating score is built from a model's internal judgement mapped onto a ten-point display. That mapping step is where the trouble starts. There is no unit behind a rating point the way there is behind a degree. The internal difference between what becomes a 6 and what becomes a 7 is not guaranteed to be the same size as the difference between an 8 and a 9, because nothing forced the mapping to be linear. What a score out of ten is actually made of covers how that internal value gets built before it ever reaches the mapping step this piece is about.

Why this matters more at the ends

Rating tools tend to compress their extremes - very few results land near 1 or 10, and a lot land in a narrow middle band. That compression is itself evidence against equal steps: if six tenths of a point near the middle separates a meaningfully different pair of results, but a full point near the top separates almost nothing, the scale is not behaving like a ruler. It is behaving like an order with some crowding.

The practical consequence is specific. An 8 is not twice a 4 in any sense that survives scrutiny - "twice as good" is a ratio claim, and ratio claims need a true zero, which this scale does not have. Even the more modest claim, that the gap from 4 to 5 equals the gap from 8 to 9, is not something the tool has promised you and usually is not true. What you can trust is the ordering: an 8 outscored a 6 on whatever this tool is measuring, on this attempt, under these conditions.

What still holds

Some tools try to reduce this problem by publishing a described rubric with named axes rather than one blended figure - Rate Cock reports six such axes rather than a total, which at least lets a reader see which specific judgement moved rather than trusting one compressed digit. Ordinal is not useless. Ranking a set of your own submissions against each other is a legitimate use of an ordinal scale, because ranking only needs the order to be right. Making a submission comparable is worth pairing with this: once the conditions are fixed, a rise from 6 to 7 across two attempts is at least evidence of movement in the right direction, even if the size of that movement is not precisely quantified.

What does not hold is arithmetic across the gap. Averaging two scores from different tools, computing a percentage improvement, or treating a half-point shift as a fixed, comparable unit of change all borrow interval-scale logic that an ordinal display has not earned. Statisticians still argue over where the line sits. Norman (2010) describes reviewers objecting that Likert-scale data are ordinal, "so parametric statistics cannot be used", and argues that studies back to the 1930s show those methods are robust anyway. That defence is about analysing many responses across a group, though, not about reading a single score as a ruler. The complete guide to comparing ratings works through which comparisons survive this problem and which do not.

The same caution applies to whatever produced the number in the first place. A model estimating a property from a photo is working from evidence with its own error bars before the score is even mapped onto ten points - the accuracy limits of that estimation step are a separate subject from the display question here, but the two compound. A human reader assigning a score from direct judgement faces a related problem from the other direction: consistent internal ranking, translated onto a scale with the same borrowed precision. And however the number was arrived at, a claim about physical size in centimetres from a tape is an interval measurement in a way a rating never will be - a different kind of number, from a different kind of instrument, and not a fair standard to hold a rating against.

Read the score as an order, not a ruler, and most of the confusion clears up on its own.

Read next

Full archive