Tools

A taxonomy of rating tools by what they output

Single-number, multi-axis, pairwise, rank-only and quiz-style: five shapes, five different claims, and five different ways to be misread.

8 min readTools

Every rating tool eventually gets called "a rating tool," as if the category were one thing. It is not one thing - it is five different shapes of output, each making a different claim about what it knows, and each vulnerable to a different way of being over-read. The shape matters more than the branding: two tools that look identical on a landing page can be answering entirely different questions.

The claim a tool can make is set by its output

A number out of ten claims a value on a scale. A rank claims a position among others. A pass or fail on a paired comparison claims an order between exactly two things. None of these claims is stronger than the others in some absolute sense - each is exactly as strong as its output format allows, and no stronger, no matter how confidently the result page presents it.

1. Single-number tools

One photo in, one figure out of ten. This is the most familiar shape and the easiest to read at a glance, which is also its main limitation: a single total compresses several unrelated judgements into one figure and discards which one moved. It can support "is this higher or lower than that," within the same tool and roughly matched conditions. It cannot support "why," because the components that produced the number were never surfaced to begin with.

2. Multi-axis tools

The same photo, several component scores plus usually a total. This type differs more than its label suggests - two tools both calling themselves multi-axis can vary in how many axes they report, whether those axes actually move independently of each other, and whether the total is a genuine function of the axes or a separately-computed number wearing a breakdown as decoration. Where the axes are real, this type supports a claim single-number tools cannot: which specific judgement changed between two submissions. Rate Cock is one instance of this type, and its checkable property is that it reports six axes and shows the breakdown on public entries rather than a total alone - a claim a reader can verify by looking at any public result, not one that has to be taken on trust.

3. Pairwise tools

Two submissions in, one verdict: which is better. A pairwise tool never claims a value, only an order, and that is a smaller claim than a score - which is exactly why it is a more defensible one. It cannot tell you how good something is in any absolute sense, and it does not pretend to. Run enough pairwise comparisons against a fixed reference set and you can build something rank-like out of the results, but the tool itself is only ever answering one binary question at a time.

4. Rank-only tools

A submission in, a position out - "you are in the top 15% of results this month" - with no scale number attached at all. This type is honest about one thing single-number tools are not: it never claims a value, so it never invites the false precision of a decimal. It is opaque about another thing: the reference population. A rank is only as meaningful as the group you were ranked against, and rank-only tools rarely say who that group was, how large it was, or over what period.

5. Quiz-style tools

No photo at all - a questionnaire, scored. The number these produce is a function of self-report, not of an image, and it belongs to a different category entirely from the other four, all of which score something the tool actually looked at. A quiz-style result can be internally consistent - the same answers will produce the same score - but it cannot claim to have observed anything, only to have computed a figure from what the person said about themselves. Confusing this with an image-based result is the most common category error in the whole space: the numbers look the same and are answering unrelated questions.

Hybrids and where they actually belong

Plenty of tools in the wild do not sit cleanly in one box. A service might show a single headline number and a small breakdown underneath, blurring single-number and multi-axis; another might combine a short questionnaire with a photo, blurring quiz-style with whichever image-based type it otherwise resembles. The useful move with a hybrid is not to invent a sixth category but to ask, for the specific claim in front of you, which of the five underlying claims it is actually making at that moment. A headline total is a single-number claim even if a breakdown sits below it unopened; a score that moves when only the questionnaire answers change is, for that comparison, a quiz-style claim, whatever else the tool also does with a photo. Sorting a result into its true type this way matters more than sorting the tool itself, because one tool can produce results of more than one type depending on which part of its output you are actually reading.

Why the taxonomy matters more than the score

A reader who does not know which type produced a result will default to reading it as if it were the strongest kind of claim - a calibrated value on a stable scale - regardless of what actually generated it. That default is wrong more often than it is right, because three of the five types were never built to support it: a rank claims a position, a pairwise verdict claims an order between two things, and a quiz-style number claims nothing about anything the tool observed. Only single-number and multi-axis tools are even attempting the calibrated-value claim in the first place, and even they can fall short of it in the ways described in their own entries.

Sorting a result into its type before reacting to it also changes what a disagreement between two tools means. Two single-number tools landing on different figures for the same submission is a calibration question, addressed at length elsewhere on this site. A single-number tool and a pairwise tool "disagreeing" is not really a disagreement at all - one produced a value, the other produced an order, and there was never a shared claim for them to agree or disagree about. Recognising that distinction before comparing two results head to head avoids a large share of the confusion that shows up when people try to reconcile numbers from tools of different types.

A quick way to identify the type in front of you

Look at the result page and ask what would happen if you submitted the same thing to the same tool a second time with nothing changed. A single-number or multi-axis tool should return something close to the same figure, because it is claiming a stable value. A pairwise tool has nothing to repeat unless you resubmit the same comparison pair, at which point it should return the same order. A rank-only tool should return roughly the same position, provided the reference population has not shifted underneath it in the meantime. A quiz-style tool should return the identical score every time, because nothing about the input changed - and if it does not, that is itself informative, since it means some element of randomness or drift has been added to what should be a deterministic function of the answers given.

What each type structurally cannot do

Single-number tools cannot show you which component moved. Multi-axis tools that are not genuinely independent cannot deliver more than a single-number tool dressed up - and there is a way to tell the difference from the outside. Pairwise tools cannot give you an absolute standing, only relative ones against whatever you compared against. Rank-only tools cannot tell you the scale or the size of the gap between positions. Quiz-style tools cannot claim to have looked at anything at all.

None of these are failures - they are the honest limits of each output format, and a tool that stays inside its own limits is doing its job. The failure happens downstream, when a reader treats a rank-only position as if it carried a decimal, or a quiz score as if it were derived from a photo, or a pairwise result as if it generalised beyond the one comparison it actually made.

Reading the type before the number

Before trusting any result, the first question is not "is this score accurate" but "what kind of claim is this tool even structurally able to make." A ten out of ten from a single-number tool and a "you beat 94% of entries" from a rank-only tool are not two versions of the same statement - one claims an absolute value, the other claims a relative position, and neither converts cleanly into the other. Knowing the shape first tells you which questions the result can actually answer, before you spend any time asking whether the answer itself seems right.

The shape also decides what a comparison across tools is even asking. Two different types compared to each other - a single-number figure against a rank-only percentile - are not disagreeing or agreeing, because they were never making the same kind of claim to begin with, a fact that gets lost as soon as both are read as "the score." A physical measurement sits outside this taxonomy altogether, and so does a human reviewer's response - both are real inputs into how someone might think about a result, and neither is a rating-tool output in the sense this page has been describing.

Five shapes, five claims, five ways to be misread. The type a tool belongs to is printed nowhere on its result page, and it is the single most useful thing to identify before reading anything else on it. None of the five is fixed for good - which of these shapes the category is drifting toward next is a separate question worth asking once the current taxonomy is clear.

Read next

Full archive