Tools

Where the gap between raters is widest

Two tools tend to agree in the middle and diverge at the ends, because the ends are where scale design differs most.

By Updated 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

Two rating tools agree most on middling submissions and disagree most on unusually high or low ones. That is not random: the ends of a scale are where each tool's own anchoring and compression choices show up, while the middle runs on shared convention.

Why the middle agrees

Most tools centre their output somewhere in a crowded band, rarely at the true midpoint of the scale. A submission that lands in that crowded band is being scored against the part of each tool's distribution that has the most data behind it and the most convention shaping it - every tool has seen thousands of ordinary submissions, and the pressure to keep the middle populated and non-alarming is similar across services. Two tools independently converging on similar conventions for "ordinary" will naturally produce similar numbers for ordinary submissions.

Why the ends pull apart

The ends of the scale are where each tool's specific design choices actually show up. How a scale is anchored at its top and bottom is one of the least standardised decisions a tool makes - some pin the maximum to a rubric ceiling that is nearly impossible to reach, some pin it to a percentile within their own population, and some do not document an anchor at all. Two tools with different anchoring will agree closely near the centre, where both anchors are far away and irrelevant, and diverge exactly where the anchor decision starts to bite.

Compression makes this worse. The top of a generous scale has to hold a disproportionate share of the population, which flattens the gaps between genuinely different results up there. A tool with heavy top compression and a tool without it will barely disagree on a 6, because neither is compressing that region, and will disagree substantially on what should be a 9, because one of them ran out of room to express it and the other did not.

A rough test you can run yourself

You do not need access to either tool's internals to see this pattern. Take a handful of your own submissions across a range of quality and run each through two tools, noting the gap between the two totals for each one. This is a rough version of what Bland and Altman (1986, The Lancet) proposed for comparing two measurement methods: study the differences between them, not their correlation, which they call "misleading". Two tools can rise and fall together across your submissions without ever agreeing on a number. If the gap stays roughly constant across the middling submissions and grows sharply for the highest and lowest ones, that is the anchoring and compression effect showing up directly in your own numbers, rather than something you have to take on faith from a general argument about scale design. It also tells you something about the two specific tools you used: a bigger gap at the extremes means their anchors are further apart than their middles suggest, even if their average behaviour on ordinary submissions looks nearly identical.

What this means for reading an extreme result

An unusually high or unusually low number is the least portable result you can get. It is the one most likely to look completely different on a second tool, not because the second tool disagrees with your submission but because the two tools built different machinery for exactly that part of the range. Comparing an extreme score across tools by converting to rank rather than reading the raw number sidesteps most of this, because rank does not depend on where either tool decided to put its ceiling.

A middling result travels better for the opposite reason: it sits in the part of both scales that was built with the least idiosyncrasy, using the most shared convention. If you only have budget to check one number across two tools, checking an extreme one tells you more about how the tools differ and less about your submission than checking a middling one would.

Where this does not apply

None of this is about whether a physical measurement disagrees at the extremes - a length taken with a tape has no scale anchor to diverge from in the first place, since it is not mapped onto a bounded rubric at all. The underlying model driving the subjective end of the score is a separate question from the scale built on top of it, and it is the scale, not the model, that is doing most of the diverging described here. A human reviewer diverges from a model-based tool for a different reason again - a person is not working from a fixed anchor at all, so their extremes and a tool's extremes are not comparable in the way two tools' extremes are. Rate Cock reports per-axis scores rather than one blended total, which at least lets a reader see which specific axis hit a ceiling rather than guessing from a single compressed number at the top of the range.

Read next

Full archive