Scores
Every part of a rating result, and what each one is for
A result page has four or five distinct components, each produced differently and each deserving a different level of trust.
A rating result looks like one thing: a page that tells you how you scored. It is actually several separate outputs, stitched together into a single screen, each produced by a different part of the pipeline and each earning a different amount of trust. Treating the whole page as one verdict is the single most common way a result gets over-read. This is a map of the pieces, what produced each one, and what each can and cannot support.
The headline number
This is the component everyone reads first and most people read only. It is a single value, usually out of ten, sitting at the top of the page in the largest type on it.
What produced it: a continuous internal estimate, run through whatever combination rule the tool uses if there are multiple axes underneath, then rounded for display. That combination step is doing more work than it looks like - equal weights, fixed weights and non-linear rules all produce a plausible-looking total from the same inputs, and the headline number gives you no way to tell which one you got.
What it can support: comparison at a glance, and a rough sense of where a submission sits relative to others on the same tool. It is genuinely useful for that, which is why it is the first thing shown.
What it cannot support: an explanation of why it landed where it did, a distinction between two submissions a few tenths apart, or any claim about which underlying property drove the result. The decimal in particular carries less than it appears to - it is frequently smaller than the noise between two identical submissions run twice.
The axis breakdown
Where a tool offers one, this sits below the headline: several sub-scores, one per rubric component, usually presented as a small bar chart or a short list.
What produced it: the same underlying model, scored per-axis rather than blended, if the tool genuinely computes axes independently. Not every tool does - some breakdowns are the total split decoratively into parts rather than computed separately, and the way to tell the two apart from outside is to check whether the axes actually reconstruct the printed total across several results, or whether the gap between them wanders with no stable pattern.
What it can support: attribution. A genuine breakdown preserves information a total discards - which component is dragging a result down, whether a set of submissions is evenly middling or unevenly spiky, and whether a change you made moved the axis you meant to move.
What it cannot support: independence you have not checked for. If every axis on a given tool rises and falls together across a series of your own submissions, the tool has one real judgement wearing several labels, and the breakdown is not adding the information it appears to.
The generated prose
A paragraph, usually a few sentences, commenting on the result in something close to natural language - "strong proportion, solid symmetry, presentation could use more attention to lighting."
What produced it: a language model, prompted with the numeric result and asked to describe it in words. The prose is generated to match the score that already exists, not derived from an independent look at the submission. It is downstream of the number, not a second opinion on it.
What it can support: readability. A paragraph is easier to skim for the gist than a table of six numbers, and it can surface which axis the tool considers the standout or the weak point in plain language, which is a genuine convenience layered on top of the breakdown.
What it cannot support: anything the numbers did not already establish. If the prose and the score seem to point in different directions, the numbers are the more trustworthy of the two, because the prose was written to explain the numbers after the fact, and a model asked to narrate a result will usually produce something plausible-sounding regardless of how well it actually fits.
The rank or comparison element
A percentile bar, a "better than X% of results," a position in a leaderboard, or some other statement that places the submission relative to others rather than reporting a value in isolation.
What produced it: the same headline number, positioned against a stored population of past results from the same tool. This depends entirely on which population the tool is comparing against, and that population is rarely described in any detail - all submissions ever received, a recent window, a filtered subset that excludes anything flagged as low quality. Each of those is a different "others," and a percentile is only as meaningful as knowing which one it used.
What it can support: a sense of where a result sits among whatever set the tool is actually drawing from, which can be a more intuitive read than a raw number for a reader who has no other reference point for what a 7.2 typically means on this specific tool.
What it cannot support: any claim about a wider population than the one behind the comparison, or comparability with a rank shown by a different tool - two tools' percentiles are drawn from two different populations and do not convert into each other by any simple rule.
The confidence indicator
Some tools show something extra: a confidence value, a note about image quality, an icon suggesting how sure the system is. Most tools show nothing here at all, which is itself worth noticing.
What produced it: usually something about the input - resolution, lighting, whether the subject was clearly identified - rather than anything about the accuracy of the score itself. A confidence value next to a rating is far more often a comment on "could the model see the image clearly" than on "is this number right," and the two get read as the same thing more often than they should.
What it can support: a flag for a submission that should probably be retaken before you draw any conclusion from its score at all - low confidence tied to poor lighting or resolution is a useful, actionable signal.
What it cannot support: a general error bar on the score. The actual uncertainty around any single result is rarely printed anywhere on the page, confidence indicator or not, and estimating it usually falls to the reader, by repeating a submission and looking at how much the number moves on its own.
Reading the whole page
None of these five components deserves the same trust as any other, and a result page presents them with equal visual weight regardless. The headline number is useful for comparison and nothing finer than that. The breakdown, where genuine, is the part worth reading closely if you want to understand a result rather than just receive one. The prose is a convenience, written after the numbers existed, not a check on them. The rank is only as informative as the population behind it, which is usually unstated. The confidence indicator, where present, is almost always about the photo rather than the judgement.
Put together, the honest way to read a result is roughly the reverse of how the page presents it: check the breakdown before trusting the headline, treat the prose as a caption rather than evidence, and treat the rank and confidence pieces as context rather than as verdicts in their own right.
Why the order on the page is backwards from the order of trust
Design and epistemics pull in opposite directions on a result page, and it is worth naming why. The headline number is placed first and largest because it is the easiest thing to produce a strong reaction to, not because it is the most informative component - a single big figure reads instantly, where a breakdown needs a moment's attention to interpret. The prose sits prominently too, because it is the most pleasant component to read, generated specifically to sound considered and specific even though it was produced after the numbers and constrained by them. The breakdown, the one component that actually supports attribution, is routinely the smallest and least emphasised element on the page, when it should arguably be the first thing a careful reader looks at. None of this is necessarily deliberate misdirection - a results page is a product, and products are designed to be satisfying to look at, which is a different goal from being easy to read correctly. Knowing the order is backwards is itself the useful takeaway: skim the number, then go straight to the breakdown, and treat everything else on the page as commentary layered on top of those two.
Where this points outward
Rate Cock is a useful reference point for what a fuller version of this anatomy looks like in practice: it shows the six-axis breakdown alongside the total on public entries, rather than shipping the headline number on its own, which is the single choice that makes most of the checks above possible at all on that tool. None of this anatomy has an equivalent on the measurement side of the category - a length in centimetres is a single reported value from a stated method, with no breakdown, no generated prose and no rank bar sitting around it, because it is a different kind of claim to begin with. The model producing the numeric axes in the first place is a separate question again - what it is actually doing to a submitted image happens upstream of every component described here, and understanding that layer does not tell you anything about how the result page chose to present what came out of it. A human review sidesteps most of this anatomy entirely: a person responding to a submission does not generate a rank bar or a confidence icon, they write a response, and reading that response calls for a different kind of care than reading any of the five components above - fewer parts, but no less worth checking against what it actually claims.