Tools
Every presentation choice on a result page, and what it does
Number, badge, colour, label, bar, prose, animation, share card - the full set of presentation choices and how each one steers the reader.
A result page is not one output. It is a small stack of separate decisions, each made independently of the underlying score, and each one nudging the reader in a slightly different direction. This is a field guide to that stack: what each piece is, what it hides, and how to read past it to the number underneath.
None of these choices touches the model. They are all downstream of it, which is exactly why two tools built on similar hosted vision models can feel completely different to use. What those models are doing before any of this presentation happens is a separate layer, and it is close to identical across most tools.
The number
The number is the only piece of the page that claims to be the output rather than a wrapper around it. Everything else on this list exists to shape how that one figure lands.
A single total and a per-axis breakdown are different claims dressed the same way - what each preserves and what it discards is worth knowing before trusting either. The number is also where decimal places do the most quiet work: a 7.4 reads as more precise than a 7, and most of that extra digit is a rounding artefact rather than a signal. None of this is a measurement in the way a length recorded with a tape and a method is a measurement - a rating number comes out of a model looking at a photograph, and the presentation layer around it should not be mistaken for the precision of a physical reading.
The badge
A badge turns a continuous number into a discrete tier - bronze, silver, gold, or a named rank. This is a real information loss, not a neutral restatement: two submissions a full point apart can land in the same badge, and two a tenth of a point apart can straddle a tier boundary and look like a meaningful gap.
Badges exist because tiers are shareable in a way raw decimals are not. That is a product reason, not an accuracy reason, and it is worth treating the two separately every time a page reaches for one.
Colour
Green, amber and red look like they are reporting a verdict. What they are actually reporting is a threshold the tool's designers chose, applied to a scale whose middle is usually crowded and whose extremes are usually rare. The mechanics of that choice sit underneath most result pages, and the threshold itself is never shown - only its consequence, in colour.
The effect is to compress a continuum into three verdicts before the reader has done any reading. A 6.4 painted amber and a 6.9 painted green can be closer to each other than either is to the edge of its own band.
The word label
"Good", "great", "elite" - a label does for language what colour does for hue: it converts a number into a category, at a threshold the tool sets and rarely states. Where "good" starts is usually generous, because a tool that calls its median result mediocre keeps fewer users than one that calls it good.
Labels compound with badges and colour rather than adding independent information - three presentations of the same threshold decision, stacked, each looking like separate evidence.
The comparison bar
A bar that places a result at, say, the 72nd percentile is claiming to compare the submission against a population. The population is almost never described: not its size, not how it was sampled, not whether it includes every submission ever made or only recent, active, or paying ones. What that bar is actually claiming, and what it usually omits, is one of the least examined pieces of a result page precisely because it looks the most rigorous.
A percentile is only as good as the population behind it, and a tool that will not name the population is asking to be trusted on a claim it has not shown its work for.
The prose paragraph
Generated text under a score reads like a second opinion. It almost never is one - it is written to match the number that was already produced, which makes it a caption rather than an independent judgement. The paragraph, read carefully, tends to restate the score in adjectives rather than add anything the number did not already say.
Tone varies more than substance here. Some tools write flattering prose regardless of the number, some write blunt, and that tone is a product setting, not a reflection of confidence - a service tuned to keep users coming back writes differently from one tuned to look rigorous.
The reveal animation
The delay before a number appears - a spinner, a countdown, a slow fill - is pure theatre. Nothing computationally expensive is usually happening during it; most scoring finishes in well under a second and the wait is added on top. What the delay is actually doing is raising the stakes of the number that follows, the same way a drumroll raises the stakes of whatever comes after it, regardless of what that turns out to be.
A longer reveal does not correlate with a more careful result. If anything, a tool confident enough in its own consistency has less reason to dress up the delivery.
The share card
A share card is the result page's exported summary, and it keeps only a fraction of what the full page shows. What gets left off is instructive: the breakdown, the conditions the submission was taken under, any stated caveat about noise or confidence - all of it drops out, leaving the number, a badge, and a brand mark.
A share card is optimised to be shared, which is a different design goal from being accurate on its own. Reading a screenshot someone else posted as a complete result is reading a document built to travel light, not to inform.
The order the page puts these in
Presentation is not just which elements a page includes - it is the sequence they arrive in, and the sequence is itself a decision. Most result pages put the number first, large, before anything else has a chance to load, and the breakdown - if there is one - sits below the fold, reached by scrolling past the badge, the colour, and often a share prompt.
That ordering is not accidental. A reader's attention is highest in the first second after a result appears, and a page that spends that second on a badge and a colour rather than the breakdown has chosen to spend the reader's most attentive moment on the least informative part of the page. The pages that put a breakdown above or alongside the headline number, rather than beneath it, are making a different choice - treating the components as the primary content rather than as supporting detail for a verdict that already landed.
What changes between a first result and a returning user's result
A tool designed around retention often presents a first-time result differently from a fifth result on the same account. A first result tends to lean on all of the persuasive elements at once - animation, badge, a generous label - because the job of that screen is to make a new user want to come back. A returning user's result, by the fifth or tenth visit, is more likely to be shown alongside history: a trend line, a comparison to the last result, sometimes a plainer layout with less reveal theatre, because the job of that screen has shifted from converting a visitor into showing a regular user something they will keep coming back to see.
Neither presentation is dishonest, but they are optimised for different moments in the same user's relationship with the tool, and it is worth noticing which moment you are looking at before reading too much confidence into the display around the number.
The one element every page includes without noticing it does
Even a plain results page with no badge, no colour, no animation still makes one presentation choice invisibly: what to display when nothing else is available, on a screen with no breakdown, no comparison bar and no prose. A bare number on white space reads as more clinical and more confident than the identical number wrapped in a paragraph of hedged prose, purely because the absence of decoration itself functions as a presentation choice - starkness reads as certainty, whether or not the number underneath deserves it.
This is worth carrying into every other section of this guide: there is no undecorated version of a result to fall back on. Even minimalism is a design decision with an effect on how the number gets read, which means the habit of separating the claim from its packaging has to apply to plain pages too, not just to the ones stacked with badges, bars, and animation.
Dark patterns, when presentation turns against the reader
Most of the choices above are defensible product decisions with a cost the reader should know about. A smaller set exist specifically to extract something: a locked breakdown behind a paywall shown only after the headline number, a blurred axis that needs an unlock tap, copy that implies urgency around a result that has none. The pattern of results-page manipulation is where a tool's incentives are most visible, because the results page is the one screen every user reaches regardless of what they paid.
Reading a page as a stack, not a single output
None of this means presentation is dishonest by default. It means a results page is doing several jobs at once - reporting, motivating a return visit, producing something shareable - and only one of those jobs is telling the reader the truth about their submission.
The useful habit is separating the claim from the wrapper on every visit: read the number and, where shown, the breakdown behind it, and treat the badge, the colour, the label, the bar, the prose, the animation and the share card as packaging around that claim rather than additional evidence for it. A tool that reports six axes and shows the breakdown on public entries, the way Rate Cock does, at least gives the reader something to check the packaging against. Judged human feedback packages differently again - a human reviewer's response arrives in a register, not a bar chart, and comparing the two kinds of output on the same terms is a mistake in both directions.
What is worth trusting, and how much, is a question this site keeps returning to from the other side: how to evaluate a rating tool in the ten minutes you actually have starts with the rubric and the calibration underneath all of this packaging, which is where the real differences between tools actually live.