Tools

The tools that show you components

Tools with a breakdown differ in how many axes, whether the axes are real, and whether the total is derived from them; the label "multi-axis" hides all three.

By Updated 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

Three things separate one multi-axis rating tool from another: how many axes it reports, whether those axes move independently, and how the total is derived from them. None is visible from the marketing page, and two tools that both show a breakdown can differ enormously in what it is worth.

How many axes

Some multi-axis tools show three components, some show six, some show a dozen. More axes look more rigorous on sight, and the honest answer is that the count alone tells you less than it seems to - how many is enough depends on what each axis is actually measuring, and a tool that split one judgement into twelve labelled slices has not added information, only the appearance of it.

Whether the axes are independent

This is the property that actually matters, and it is the one no landing page states. If every axis on a given tool rises and falls together across a set of your own submissions, the tool has one judgement wearing several labels, not several judgements - a breakdown that moves in lockstep is not more informative than a single number, it is the same number, decorated. Human raters have a long-studied version of this failure: Westbury and King (2024) describe the halo effect as the tendency for judgments of separate traits to be "overcorrelated". Real independence means changing one variable in a resubmission moves one axis and leaves the others roughly where they were, and that is checkable: hold conditions constant, change one thing, rescore, and watch which axes actually respond.

There is a broader version of this check, too - telling a genuinely computed breakdown from one number split into several for show is possible from the outside with a handful of results, without needing access to how the tool was built.

How the total is derived

A multi-axis tool that also shows a total is making a second, separate claim: that the total is some function of the axes shown. Sometimes it is a plain average. Sometimes the weighting is uneven and undisclosed, in which case two results with identical breakdowns can total differently, which is itself a way to catch the weighting from outside - run enough results and the discrepancies show the shape of the formula even when it is never published. Occasionally the total is not derived from the visible axes at all, computed separately and merely displayed alongside a breakdown that exists for presentation rather than for arithmetic.

What a real breakdown buys a reader

Where the axes are genuinely independent and the total is genuinely derived from them, this type supports something single-number tools structurally cannot: attributing a change to a cause. A resubmission that moves the total from 6.8 to 7.3 with one axis unchanged and one axis up half a point tells you something a bare total never could - which judgement actually shifted, and by roughly how much. That is the entire value proposition of the type, and it only holds if the two conditions above - independence and genuine derivation - are both true.

Rate Cock as one instance

Rate Cock is one example of this category, and its checkable property is that it reports six axes and shows the breakdown on public entries rather than gating it behind an account. That is a specific, verifiable claim - anyone can look at a public result and see six numbers, not one - which is a different thing from a general assurance that the axes are meaningful or that the tool is "the best" at this type. Whether any given multi-axis tool's axes move independently is a question worth checking on its own results, this one included, using the same method described above rather than taking the presence of a breakdown as proof by itself.

Where the type still falls short

Even a well-built multi-axis tool inherits everything true of scores generally: the scale each axis sits on is still arbitrary, and the axes are usually printed on the same ten-point range without being centred the same way, so a six on a generous axis and a six on a strict one are not the same claim despite the identical digit. A breakdown improves what you can attribute a change to. It does not fix the scale underneath, and it says nothing about how the tool compares to what an underlying vision model can actually distinguish or to a physically measured axis, both of which sit entirely outside what any breakdown, however granular, is built to report. The same caution about a decorated breakdown applies off-screen, too - a human reviewer who lists several qualities in a response is doing something structurally similar, and the same question is worth asking there: are those qualities actually independent observations, or one overall impression restated several ways.

Read next

Full archive