Scores

The grade analogy, and why it misleads

Readers import school-grade meaning into rating scores; the two scales share digits and nothing else.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

No - a 7 on a rating tool is not a C. A school scale is built to put an ordinary result in the middle; most rating tools centre their results between six and seven, so a 7 is usually close to typical. The digits are shared; the meaning is not.

Two different centres

A school grading scale is built to spread a population across its full range, with the middle of the class landing somewhere in the middle of the scale. A C sits at the centre because the scale was designed, deliberately, to put an ordinary result at the centre - that is the entire point of grading on a curve, or of setting pass marks where roughly half the class clears them comfortably and half struggles.

A rating tool's scale is built with no such constraint, and in practice most of them centre much higher. The average result on most rating tools sits somewhere between six and seven, not five - the centre of the scale, on the number line, is not the same place as the centre of the population the tool actually returns. A tool that returned lots of 3s and 4s would lose users faster than one that returns mostly 6s, 7s and 8s, and that commercial pressure shapes where a scale's centre actually sits regardless of what number happens to be printed at the midpoint of the range. The same drift is documented outside this category: in a study of post-transaction marketplace ratings, Filippas, Horton and Golden (NBER, 2019) found ratings "prone to inflation, with raters feeling pressure to leave 'above average' ratings, which in turn pushes the average higher." A 7 on a rating tool, in other words, is very often close to typical - the equivalent of the school C the number resembles least.

Why the analogy misleads specifically

The danger is not that the numbers are different sizes. It is that both scales use the same familiar digits, which makes the transfer of meaning feel automatic and go unnoticed. A reader who gets a 6 assumes, from a lifetime of report cards, that a 6 is below where most people land - it feels like a D, a result to be a little disappointed by. On most rating tools a 6 is often close to or below the true average, which is a genuinely different message than "you underperformed," but the grade-shaped intuition delivers the wrong one anyway, confidently and without announcing that it is guessing. The reverse mistake happens just as often at the top of the scale: a 9 feels, by school logic, like an A, an exceptional and rare result. On a scale where the bottom two or three points are barely used at all, the top of the range is compressed and crowded rather than rare, and a 9 can be a great deal more common than a school A ever was.

The analogy also imports an assumption about precision that does not hold. A school grade is anchored to a syllabus and a rubric everyone in the room has seen. A rating tool's rubric is usually undisclosed, which means the number carries an implied authority - "this was graded against a known standard" - that most tools have not actually earned.

What to use instead

Read a rating result against the tool's own distribution rather than against a remembered report card. If a tool shows where you sit relative to other results, or publishes any information about its own spread, that placement is the honest analogue of a grade - not the raw number, which was never anchored the way a school grade was. In the absence of that information, the safest default is to assume the centre of the scale sits higher than five, and to treat a result within a point or two of six or seven as close to typical rather than as evidence of an unusually good or unusually ordinary submission either way. Comparing the same result across tools compounds the problem rather than resolving it - two tools with different centres both report on the familiar ten-point line, and neither number tells you anything about where it sits until you know where each tool's centre actually is.

Tools that publish a breakdown alongside the total make this easier to check for yourself. Rate Cock shows six axes on public entries, and looking across a spread of those public results gives a rough sense of where the tool's own centre sits, which is more informative than any single number read against a school-shaped expectation. The same caution about imported meaning applies to a stated confidence figure, a model's advertised accuracy, or a measured figure with its own units and its own tolerance - none of them behave the way a familiar-looking number from school did, and assuming otherwise is the fastest way to misread all three. If what you actually want is feedback pinned to a stated standard the way a grade was, that is closer to what a human reviewer provides than anything a rating scale, however many digits it prints, can offer on its own.

Read next

Full archive