Scores

What the written feedback in a result actually is

The prose is generated to match the number, not derived from an independent look; treat it as a caption, not a second opinion.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

The paragraph under a rating score is not a second opinion. It is generated to be consistent with the score that was already produced, which makes it a caption rather than a verdict, even though it reads like a second look written by something that studied the submission independently.

Where the text actually comes from

The score is computed first. The prose is filled in afterward, conditioned on that score and on whatever axis values the tool already has. A result of 8.1 with a strong symmetry axis gets language that praises symmetry; a result of 6.4 gets language calibrated to sound encouraging without contradicting a middling number. This is not a secret and it is not a scandal. It is how a language layer sits on top of a scoring layer in almost every tool that produces both, the same underlying model doing the scoring that does the rest of the assessment work before any text gets generated at all. The five things that actually separate rating tools covers rubric and calibration; the prose is downstream of both and adds nothing to either.

Why it sounds so confident

Generated text tends toward specificity because specificity reads as authority. A sentence naming a feature ("strong proportion, slightly softer definition") sounds like it came from close inspection. Often it is closer to a template with the axis values dropped in. That is not automatically worthless - a template built from real per-axis numbers is reporting something true, just not independently derived. Research on language models' own explanations shows why fluency is not evidence: Turpin and colleagues (2023) nudged models toward wrong answers with biasing features in the prompt, saw accuracy drop by as much as 36% across 13 tasks, and found the explanations rationalised those answers without mentioning the bias - text that, in their words, "can be plausible yet misleading." The failure mode is a template that would say roughly the same thing regardless of the photo, decorated with a number to make it feel bespoke.

What to actually read for

Two tests separate a useful paragraph from decoration.

Does it name something that could have gone the other way? A sentence that only a genuinely lower-scoring submission would not receive is doing work. A sentence that fits almost any result in the tool's population is filler.

Does it match the breakdown, if there is one? If the tool publishes a per-axis breakdown alongside the prose, check that the paragraph's claims track the axis it is praising or hedging on. The full anatomy of a result page breaks the number, the breakdown, the prose and the rank into four separately-trustworthy components, and the prose is consistently the least independent of the four. Rate Cock is one tool that publishes that kind of per-axis breakdown alongside its written result, which is what makes this particular check possible rather than theoretical.

Generic encouragement - "great submission," "you're doing well" - carries no information and should be discounted entirely regardless of how it is phrased. It is the part of the result written for retention, not for accuracy.

The comparison worth making

If a tool shows both a breakdown and a paragraph, the paragraph is redundant with something you can check more directly: the axis values themselves. Treat the prose as a reading aid for the breakdown you already have, not a fifth data point. On tools that show only a total and a paragraph, the paragraph is the only qualitative information available, which raises the stakes on distinguishing specific claims from generic ones - and lowers the amount of trust either deserves on its own. Whether a tool leans on a number badge or a paragraph as its primary output is itself a design choice worth noticing, since the two formats invite different amounts of scrutiny from the same underlying score.

The same distinction shows up outside automated tools. A human reviewer also writes prose in response to a submission, but it is generated by a person reading the actual material rather than by a model conditioned on a number it already produced - a genuinely different kind of text, even when the two read similarly on the page. That's a different product with a different set of tradeoffs, not a better or worse version of the paragraph under an automated score. None of this touches how the underlying photo was captured or processed before it was measured at all - the prose is entirely a scoring-side artefact, generated after the image has already done whatever work it was going to do.

What the paragraph cannot do, on any tool, is function as an independent second opinion. It was never given the chance to disagree with the number. Read it for the specifics it names, discount the rest, and go to the breakdown when you want to know what actually moved.

Read next

Full archive