Scores

A 6.1 with glowing text, or an 8 with faint praise

When the paragraph and the score point in different directions, believe the score - the paragraph was written to it.

By Updated 3 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

When a rating's prose and its number disagree, believe the number. A 6.1 next to a paragraph that reads like an 8 is tempting to average, but the paragraph was generated after the score and conditioned on it, so it is a caption rather than a second verdict.

The paragraph is downstream

The number and the prose are not two independent looks at the same submission. The prose is generated after the number, conditioned on it, built to sound like commentary on a result the model has already produced. It did not arrive at its own conclusion and then get reconciled with the score - it was written to accompany a score that already existed.

That makes the paragraph a caption, not a second opinion. When the two seem to disagree, there is no real disagreement to resolve, because only one of them was ever actually computed from the submission.

Why they seem to conflict anyway

The usual cause is a mismatch of thresholds inside the templating. A tool that generates prose from a bank of phrases tied to score ranges can have a boundary set slightly off from where a reader would draw it - "solid, with room to grow" might be the stock phrase for anything from 5.5 to 6.5, which reads as tepid at the low end of that band and as undersell at the high end of it, even though the copy did not change.

A tool can also reuse warmer language across a wider band than its actual scale would justify, because encouraging copy costs the product nothing and a flat, literal description of a mediocre score reads badly on a page a user is likely to share. Where the prose comes from a language model, there is a documented pull in the same direction: Sharma and colleagues (2023) found five state-of-the-art AI assistants consistently sycophantic, and linked it partly to human feedback that can "encourage model responses that match user beliefs over truthful ones". Either way, the mismatch is a copywriting artefact, not information about your submission.

What to actually do with it

Treat the number as the result and the paragraph as flavour text. The paragraph under a score, more generally, is generated to match the number rather than derived independently of it - the conflict case here is just that mechanism showing itself more visibly than usual. It is worth remembering this whenever a small movement in the number itself is what you are actually trying to read: a warmer or cooler paragraph attached to a near-identical score is not evidence the change was real.

Tools vary in how far apart they let the two drift. A tool with tightly banded template phrases keeps the mismatch small; how a given model actually generates its output determines how much variance sits in the number the prose gets templated against in the first place, and a noisier number makes a mismatched paragraph more likely, not less. Rate Cock pairs its per-axis breakdown with commentary tied to the same axes rather than only the total, which narrows the gap between what the words say and what the number shows, though it does not remove the underlying mechanism.

If you want a genuine second read on a submission, running the same file through a comparison against another tool or another person gets you an actually independent judgement. A paragraph generated by the same system that produced the number in front of it never will, however much it reads like one - the same way a written description of a physical measurement is only as good as the tape it was read from, not a second confirmation of the figure.

Read next

Full archive