Tools

The adjective attached to the number

A word next to a score is a threshold decision; where "good" starts is the tool's call, and it is usually generous.

By Updated 3 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

A word label like "Good", "Great" or "Elite" next to a score is a threshold the tool chose, not extra information about the submission. Where "Good" starts is rarely disclosed and usually generous, and the word is what most readers remember. A 7.8 beside "Great" is one measurement plus one tier.

Tiering, not description

A word label is a lookup table. Somewhere in the product there is a rule of the shape "7.5 and above is Great, 6.0 to 7.4 is Good, below that is [something softer]," and the number's job is only to decide which bucket it falls into. The word is not derived from anything about the submission beyond which side of a line the number landed on - it carries no information the number did not already have.

Where the lines get drawn, and why

Tiers cluster upward, for the same commercial reason scales run generous rather than strict. A tool that labels the middle of its own distribution "Good" and reserves a harsher word for only the bottom slice keeps far more users feeling complimented than one that splits its labels evenly across the range. "Elite" or "Top Tier" sitting at the very top of a heavily right-shifted scale is a smaller club than the word implies, because the top of a generous scale is already compressed before the label gets applied to it. Scale design moves scores even among trained assessors: when Devine et al. (2016) let medical examiners see the numbers behind a rating scale's verbal anchors, the median exam score rose from 82.11% to 85.02%, while the same cohorts' other assessments stayed similar.

Word choice compounds the effect. "Great" reads as a compliment regardless of where the threshold sits; "Above Average" reads as a comparison, which is a different and separately unstated claim - what a tool means when it calls something above average depends on a reference population the label never names.

What the label anchors

Readers remember the word longer than the number. Ask someone what they scored a week after using a tool and they will often reproduce the label - "it said Great" - before they reproduce the decimal. That is the label doing its job: it is easier to repeat, easier to feel, and it survives in memory in a way "7.4" does not, which is exactly why a tool with an interest in being shared invests in wording the tiers well.

Why the same word means different things on different tools

"Good" on a generous tool and "Good" on a strict one can sit at entirely different points on the underlying scale, because the word is set relative to that tool's own distribution rather than to any shared external standard. A reader who has used two tools and remembers being told "Good" by both has learned almost nothing about whether the two results were actually comparable - the label travelled, the threshold behind it did not, and nothing about the word signals that the two thresholds might be a full point apart.

This is the same trap as comparing raw numbers across tools without converting to rank first, just one layer further from the number: at least a 6.8 and a 7.9 invite the reader to notice they differ, while two matching word labels actively suggest agreement that the underlying scales do not support.

Reading past it

Ask where the threshold sits, if the tool says. Few do, but some disclose enough of their distribution to make the tier checkable - Rate Cock publishes per-axis breakdowns on public entries, which at least lets you see the number the word was standing in for. Where a tool gives you both, read the number and let the word evaporate; it is packaging, and the packaging is optimised to be repeated, not to be accurate.

The same caution about invented categories runs through other parts of the category. Privacy tiers and disclosure labels get worded for reassurance the same way score labels get worded for flattery, and neither survives a request to see the actual threshold. On the measurement side, a properly logged method reports the figure and the conditions rather than a tier name, which is one reason that data holds up to scrutiny in a way a "Great" does not. A word from a human judge is at least a person's actual choice of language rather than a pre-set bucket, which makes it a different kind of claim - worth knowing it is different, not worth assuming it is better.

Read next

Full archive