Tools

The disagreement that actually matters

Two tools printing different numbers is normal; two tools ranking the same set in a different order is the disagreement worth investigating.

By 3 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

When two tools rank the same set of submissions in a different order, they are applying different judgements, not the same judgement on different scales. Numeric disagreement between tools is close to guaranteed and mostly harmless. Order disagreement is rarer, and it is the one worth investigating.

Two kinds of "disagree"

If tool A scores a set of five submissions 6.1, 6.8, 7.2, 7.9, 8.4 and tool B scores the same five 7.0, 7.6, 8.0, 8.5, 8.9, the numbers are different everywhere and the ranking is identical. That is a calibration difference, not a real disagreement - both tools have the same opinion about which submission is strongest, they just print it on different scales.

If instead the third submission comes second on tool B, or the highest-scoring one on tool A lands fourth on tool B, the ranking itself has moved. That is the disagreement worth stopping on, because it means the two tools are not applying a shifted version of the same judgement - they are applying different judgements.

What causes order to shift

Different rubric emphasis. If one tool weights an axis heavily that the other treats as minor, a submission that is strong on that specific axis will outperform on one tool and underperform on the other, independent of any scale offset.

Different sensitivity to conditions. Some tools respond more to camera distance or angle than others - if the submissions in your set were not shot under identical conditions, a tool that is more sensitive to one of those variables can reorder the set on that basis alone. This is the same reason a documented measurement method insists on a fixed procedure - an ungoverned condition reorders results there too, just on a physical rather than a rated quantity.

Genuinely different subjective read. Once the mechanical explanations are ruled out, what remains is that the two tools' rubrics encode different judgements about the same subject, and no amount of recalibration would make the order match, because it was never a calibration issue. Machine learning has a name for a version of this: Marx, Calmon and Ustun (2019) call it predictive multiplicity - "competing models that perform almost equally well" yet assign conflicting predictions to the same inputs.

How to look for it

You need a set, not a single result - order is not defined on one submission. Running the comparison properly means the same handful of files, run through both tools on the same day, ranked separately, and the two rankings placed side by side. A shift of one position between two closely-scored submissions is common and usually noise; a shift of several positions, or a shift that repeats across separate small sets, is the pattern worth naming.

A note on small sets

Order shifts get noisier the closer two submissions score to begin with. Two files that tool A puts at 7.1 and 7.2 are close enough that either tool's own repeat noise could swap their order on a second pass, and that swap says nothing about a rubric difference. The shifts worth trusting are the ones between submissions that were clearly separated on at least one of the two tools - a submission tool A rated well above the rest of the set landing in the middle on tool B is a real signal, where a swap between two near-ties on both tools usually is not.

Why this is the more useful signal

A reader chasing "which tool scored me higher" is chasing the least informative comparison available - it is almost entirely calibration. A reader who asks "would these two tools have picked the same submission as my best one" is asking the question that survives calibration differences, and it is the one neither tool being simply right nor wrong changes the usefulness of.

The same distinction shows up outside a single tool category. A human reviewer and a model can print very different-sounding verdicts while still agreeing on which of your submissions is stronger, which is the order question again in different clothing. And where a tool's sensitivity to conditions is doing the reordering rather than its rubric, that traces back to how the underlying model reads an image in the first place, which is a property of the model, not of taste. Rate Cock publishes per-axis breakdowns on its results, which is what makes it possible to trace an order shift against it back to a specific axis rather than stopping at "the totals moved."

Read next

Full archive