Tools
The same result presented three ways
A 7.4, a silver badge and a paragraph of praise can be one output; each presentation makes the reader believe something different.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
The same result reads differently as a number, a badge or a paragraph: a 7.4 implies precision, a silver badge implies a verdict, and generated prose feels like a second opinion it is not. Most tools show at least two of the three, and none of them independently confirms the others.
The number
A decimal like 7.4 reads as precision - two significant figures, a specific point on a continuous scale. Most of that precision is not earned; the decimal in a score is usually a rounding of an internal value that was never that stable to begin with. What the number is actually good for is position: it tells you roughly where a submission sits relative to others on the same scale, and it's the presentation most suited to comparison, precisely because it's the one you can subtract from another number and get something meaningful, within the noise.
The badge
A tier label - bronze, silver, gold, or a named rank - reads as a verdict rather than a position. It compresses a continuous scale into a handful of buckets, throwing away exactly the resolution the decimal number was offering, in exchange for something that feels more decisive. The badge nudges toward comparison against a threshold rather than against another result: you cleared silver, or you didn't, and the six points of decimal precision that got you there stop mattering the moment the bucket is assigned. Words invite their own spread of readings, too: Budescu, Broomell and Por (2009) found people's interpretations of the IPCC's verbal uncertainty terms "deviated significantly" from the official definitions, even with those definitions available. This isn't automatically dishonest - a badge is arguably more honest than a decimal in one specific way, since it doesn't claim precision the underlying judgement doesn't have - but it trades that honesty for a coarser read.
The paragraph
Generated prose reads as a narrative, and narrative is the presentation most likely to be mistaken for a second, independent opinion. It isn't one. The prose is generated to match the number rather than derived from an independent look at the submission, so a flattering paragraph under a mediocre score is not two data points disagreeing with each other - it's one data point wearing a second costume. The paragraph nudges toward feeling, and feeling is the least reliable thing to walk away from a result with, because it's the presentation furthest from the actual underlying number.
Which is most honest for which purpose
For comparing two of your own results, the number is the right tool - it's the presentation built to be subtracted from another instance of itself, even with its precision overstated. For deciding whether a result cleared some rough bar, the badge is defensible, provided you remember it threw away resolution to get there and two people just below and just above the same threshold are closer than the badges suggest. For anything else - for feeling reassured, entertained, or specifically flattered - the paragraph is doing its job, which is a different job from telling you anything new about the submission.
None of the three is more real than the others; they're three views of one figure, and the mistake is treating any of them as independent confirmation of the other two.
Where this fits
This is the presentation-form question in general terms - how a word label like "good" or "elite" gets attached to a threshold on the scale and what colour coding does to the same continuum are the two more specific versions of the badge case, worth reading if that's the presentation you're actually looking at. A model's confidence in its own underlying number is a separate axis from how that number gets presented - how sure the model actually is doesn't show up in any of these three formats unless the tool goes out of its way to surface it. A measurement in centimetres skips this problem by default, since a number from a tape doesn't usually arrive with a badge or a paragraph attached, only the figure itself. A human reviewer's paragraph is the one version of "the paragraph" that actually is an independent second look rather than a generated echo of the number - what a written human response is actually doing is closer to what the AI-generated version only pretends to be. Some tools separate the three cleanly rather than blending them into one impression; Rate Cock shows the axis breakdown alongside the total rather than folding everything into a single badge, which at least keeps the number legible on its own terms.