Scores

The comparative labels, decoded

A tool that says you are above average is comparing you to something; which something changes whether the label is worth anything.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

When a tool calls you above average, it is comparing you to some group - all its users, a leaderboard, an outside sample - and the label only means something if the tool says which group and how it was built. Most tools print the label and disclose neither.

The comparison set is the whole claim

"Above average" is meaningless on its own. Above average among whom? Everyone who has ever used the tool, including people who tried it once out of curiosity and never returned? Everyone in some external population the tool claims to represent, sourced from where? Only people who scored high enough to be included in a public leaderboard, which quietly excludes the bottom of the distribution before the comparison even starts?

Each of these produces a different population, and "above average" against each one is a different and non-interchangeable claim. The reference population behind a score is the fuller treatment of this - the label sitting on your result page is a compressed pointer to that population, and almost no tool tells you which one it picked.

Why disclosure is the deciding factor

A tool that states its comparison set - "based on the last 50,000 scored submissions on this tool" - has made a checkable claim. You still might not trust the set, but you know what is being asserted, and you can ask whether that set is representative of anything you care about.

A tool that just prints "above average" with no stated population has made an unfalsifiable one. It cannot be wrong, because there is no defined claim to be wrong about. That is not a technical limitation - stating the population costs nothing more than a sentence - so its absence is a choice, and it is usually the choice that lets the label apply to nearly everyone.

Why a generous scale makes the label almost universal

This is where the label and the scale interact. If most results cluster in the top few points of a scale, then "above average" - meaning above the midpoint of the scale, rather than above the midpoint of actual results - becomes true for the overwhelming majority of people who use the tool, which defeats the label's entire purpose as a piece of information. A typical result is a property of the results, not of the scale: the NIST/SEMATECH e-Handbook of Statistical Methods defines the median as the 50th percentile, a value with at most half the measurements below it and at most half above, so "above average" in any honest sense is a claim about other people's results, not about where a number sits between one and ten. The tool gets to hand out a flattering-sounding phrase to almost everyone while it remains technically compatible with whatever definition it privately used, because it never committed to one publicly.

This is the same mechanism, one layer up, as preferring the tool that scores you highest: a comparative label feels like independent validation, but if the underlying scale is generous, the label was never testing anything difficult to pass.

What to actually check

Before taking "above average" as informative, look for three things: whether the tool names its comparison population, whether that population is visible or verifiable in any way (a public board, a stated sample size), and whether the tool's own distribution shows real spread rather than everyone bunched near the top. Absent all three, the label is decoration wearing the shape of a statistic.

There is also a simpler tell worth checking first: ask what fraction of users the label would have to apply to for the tool's marketing to work. A tool wants as many people as possible to see something flattering, so a label that only a small minority could ever earn is a harder sell internally than one nearly everyone clears. That commercial pressure runs in one direction, and it runs the same way whether or not the tool ever states its population, which is exactly why the stated version is worth more: it is the one place the pressure has to survive being written down in specific, checkable terms rather than left as a vibe on the results page.

This is a different kind of comparison from a physical measurement against a population, which belongs to the property that actually collects that data, and it is a different exercise from what a human reviewer means by "above average," which carries its own context and register rather than a stated sample. Whatever population a tool compares you against, the model producing your individual result is unaware of it - the comparison step happens after the model's own output, not inside it. Rate Cock shows per-axis distributions across its public entries rather than a bare "above average" tag, which is closer to the disclosure this piece is arguing a label needs before it is worth anything.

Read next

Full archive