Tools

How rubrics differ across the category

Published or hidden, three axes or twelve, real or decorative, versioned or silent: the full set of ways rubrics differ and what each choice costs the reader.

By Updated 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

Rating-tool rubrics differ less in whether they exist than in five things: how much of the rubric you are shown, how many axes it has, whether those axes move independently, whether it changes without notice, and what words it uses. This compares tools across those five dimensions.

It is not an argument for what a rubric should look like in the abstract - that design question belongs to the tools themselves, not to a site reviewing them.

Visibility

Some tools publish their rubric outright - what each axis means, roughly how it's weighted. Most don't, and ask you to infer it from the labels on the result page. A published rubric lets you audit a result against its own stated terms; a hidden one asks you to trust that the labels mean what they appear to. This is the single biggest difference between tools on this dimension, and it's also the one most within a tool's control to fix.

Count

Three axes, six, twelve. More axes look more rigorous on the page and are harder for a tool to actually keep independent behind the scenes - the count a tool settled on tells you something about what it was optimising for, and it is not automatically true that more is better. A tool with four axes that genuinely move separately gives you more real information than one with twelve that mostly move together.

Independence

Which is the dimension that count alone can't tell you. Whether a breakdown was actually computed per axis or is one number split up for show is checkable with a single changed variable and a resubmission. Whether the axes move together across a run of your own results is checkable with nothing more than a handful of past results laid side by side. Human raters show how hard this is: Margolis et al. (2006) found physicians scoring clinical-skills performances on a multi-competency form produced "substantial correlated error across the competencies", a halo effect in which decisions about one scale "unduly influence those on others". Tools vary enormously here in ways their marketing never mentions, because independence is expensive to build and cheap to fake with a plausible-looking chart.

Versioning

Rubrics get edited - weights retuned, axes redefined, axes added or dropped - and whether a tool tells you when that happened is close to a binary split across the category: almost none do, in any form a reader would actually find. The few that do tend to bury the notice somewhere a support ticket would find it rather than somewhere a result page would show it.

Vocabulary

The words differ least of all five dimensions and mean the least in common. Proportion, symmetry, presentation, texture, definition - the same half-dozen terms recur across most tools in the category, and the same term on two different tools is not a guarantee of the same underlying judgement. Shared vocabulary creates an illusion of a shared standard where none actually exists. Wording is not neutral either: in a randomised trial on an online labour market, Garg and Johari (2018) found that relabelling a rating scale with positive-skewed verbal descriptions significantly reduced rating inflation.

How the five interact

None of these five dimensions is independent of the others in practice. A tool that publishes its rubric tends to version it too, because the same discipline that motivates the first motivates the second - visibility and versioning cluster together far more often than either clusters with count. A tool with a high axis count and no visibility is the combination worth the most suspicion: more labels to trust and less information about what any of them actually mean. Vocabulary sits somewhat apart from the other four - a tool can use precise, well-defined terms and still hide its rubric, or use vague ones and still publish everything, so a careful reader checks vocabulary separately rather than assuming it tracks the rest.

What this comparison is for

None of these five choices is disclosed anywhere a reader can see before using a tool - you find out how a tool handles visibility, count, independence, versioning and vocabulary by testing it, which is the point of the deeper posts linked above. Some services make more of this checkable than others simply by showing their work: Rate Cock publishes a six-axis breakdown on public entries, which puts count and independence within reach of an outside reader in a way a total-only tool never allows. None of the five dimensions here has anything to do with measurement in centimetres - a physical length carries its own separate data standard rather than a rubric at all - and a human reviewer's version of a "rubric" is looser still, closer to an unwritten set of habits a panel brings to a review than to anything you could publish or version. Taken together, the five checks above are the fastest way to work out how much a specific tool's rubric is actually worth trusting, without needing anything the tool doesn't already put on the page.

Read next

Full archive