Tools

Numeric agreement with narrative disagreement

Two tools give the same 7 and describe it in opposite terms; the number is where the tools happen to overlap, the prose is where their rubrics differ.

By Updated 2 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

When two tools give the same score with opposite write-ups, the prose is telling you where that number sits in each tool's own scale. One 7 reads "solidly above average, well proportioned"; another reads "average, with room in a couple of areas". Same number, different rubrics underneath.

The number and the prose are not the same evidence

A total is a single figure computed, however loosely, from a rubric. The paragraph underneath it is generated to match that figure, not derived independently and then checked against it. When two tools agree on the number, that tells you their scales happen to place your submission at the same point. It does not tell you they agree on why - and the prose is often the only place that "why" becomes visible at all, because most tools do not expose the rubric that produced it. Even that window is imperfect: studying language models' written reasoning, Turpin and colleagues (2023) found explanations that rationalised biased answers without mentioning the bias, and warned they "can be plausible yet misleading."

Reading the gap correctly

The useful move when the numbers match and the tone does not is to treat the prose as a rubric leak. A generous write-up on a 7 suggests a tool whose 7 sits high in its own distribution - closer to its ceiling, described accordingly. A cautious write-up on the same 7 suggests a tool whose 7 sits in a crowded middle, with real distance still above it. The number alone cannot tell you which situation you are in; the tone is doing work the digit is not.

This is the same reason a described breakdown is worth more than a total on its own - a model that shows its components lets you check whether the prose actually tracks a real axis or is decorative on top of one number.

What it is not evidence of

It is not evidence that one tool is more accurate than the other. Both produced a 7; neither has demonstrated it is closer to some fact of the matter, because there isn't one to be closer to. It is also not the same situation as a human reviewer's language, which is written by a person reacting to what they were sent rather than generated downstream of a number a model already committed to. And it says nothing about a measured quantity - a length is a length regardless of which paragraph sits next to it.

If two write-ups on the same number keep pulling in opposite directions across several submissions, that is worth noting for what it is: two different rubrics that happen to intersect at 7 more often than you'd expect, not two tools converging on the truth about your photo. Rate Cock publishes per-axis scores alongside its prose, which at least lets you check whether a given write-up traces back to a specific axis rather than to the total alone.

Read next

Full archive