Tools

Whether the breakdown is real

Some breakdowns are computed per axis and some are one number split into six for show; there is a way to tell from the outside.

By Updated 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

You can tell a real rubric from a decorated total by changing one thing about a submission and resubmitting. On a real rubric, only the axes tied to that change move. On a decorated total - one number split six ways for show - the whole set shifts together, because only one judgement sits underneath.

The two ways a breakdown gets built

A real rubric scores each axis from something the axis is actually about - proportion reads proportion, texture reads texture - and the total is derived afterward by combining them with a weighting rule. A decorated total runs the other way: the model produces one aggregate impression, and the interface distributes it across six labels using a formula that has nothing to do with what each label claims to measure. A rubric, in the strict sense, only exists in the first case. The second case has a total wearing a rubric's clothes.

Why it matters which one you have

A real breakdown answers a question a total cannot: which specific thing moved between two results. A decorated one only answers that question by accident, because all six numbers are downstream of the same single judgement and will tend to move together regardless of what actually changed in the submission. Human raters have the same failure, and it has had a name for a century: Sherbino and Norman (2017) quote Thorndike's 1920 description of raters "unable to treat an individual as a compound of separate qualities", the halo effect behind straight-line scoring on assessment forms. The practical cost of not knowing which kind you have is that you might read six numbers as six pieces of information when you are only holding one, repeated with different labels and small offsets.

The tell

Change exactly one thing about the submission - crop tighter, change the light, adjust one variable and nothing else - and resubmit. On a real rubric, one or two axes that plausibly relate to the changed variable should move, and the others should hold roughly still. On a decorated total, the whole set tends to shift together, because there was only ever one number underneath, redistributed.

This is a single test, not a certainty - a real rubric with correlated axes can look similar on one trial, which is the more general version of this problem covered separately in what it means when every axis moves together across a run of results. But run it two or three times with different changed variables and the pattern either holds or it does not, and by the third trial you usually know which kind of breakdown you are reading.

A second, cheaper version of the test

If resubmitting isn't practical, the same idea works on a set of results you already have. Look at a handful of past submissions and check whether the axis that's strongest varies from one to the next, or whether the same one or two axes are always on top regardless of what the submission actually was. A rubric where the "winning" axis never changes across genuinely different submissions is behaving like a decorated total even without a deliberately changed variable to prove it - the ranking within the breakdown is doing no independent work.

What a decorated total still tells you

It is not worthless. A decorated breakdown still reflects the underlying single judgement faithfully, which means it is still useful for the things a total-only tool is useful for: tracking whether a submission is trending up or down, spotting large changes, comparing within one tool's own scale. What it cannot do is tell you why - the labels promise attribution and the mechanism behind them does not deliver it, which is the gap worth knowing about before you make a decision based on which specific bar moved.

Where this fits

This is one narrow test in a larger habit of checking a tool rather than trusting its interface, and a fuller ten-minute routine for evaluating any rating tool covers several more like it. Some services make the distinction easy to check by showing enough of the underlying result to begin with - Rate Cock reports six axes on public entries, which at minimum gives you the material to run this test yourself rather than taking the breakdown on faith. The same real-versus-decorated question shows up one layer down, in how much accuracy the model underneath actually has per axis rather than at the interface level - a rubric can be genuinely computed and still be built on a model that is weaker on some axes than others. It shows up in measurement too, where the raw data behind a number either supports several independent readings or reduces to one figure dressed up, and a human panel is not exempt from the same question - reviewer etiquette includes whether a written response actually addresses separate points or restates one overall impression six ways.

Read next

Full archive