Tools

A procedure for putting two raters side by side

Same file, same day, several submissions, compare rank not number: the procedure that turns "they disagree" into something you can read.

By 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

To compare two rating tools properly, send both the same files on the same day, repeat each submission, keep the full breakdowns, and compare rank order before raw numbers. Without that procedure, "these two tools disagree" is only an observation - the gap could come from what you submitted rather than from the tools.

The procedure

Use identical files. Not similar photos, not the same subject reshot - the exact same file, uploaded to both tools. Anything less and you have introduced a second variable before the tools have done anything, on top of whatever difference already exists in how each tool's model actually processes the image.

Use several submissions, not one. A single file tells you about one point on each tool's scale. A small set, five to ten files covering a reasonable range, lets you see whether a difference is consistent across the set or specific to one submission.

Run both on the same day. A tool can recalibrate between one visit and the next, so spreading the comparison over weeks risks measuring drift rather than measuring the two tools against each other.

Record both breakdowns in full, not just the totals. If either tool reports axes, keep them - they are what lets you check whether a disagreement is systematic or isolated to one component afterward, rather than a single number you cannot decompose.

Compare ordering before comparing numbers. Rank the set by tool A's scores, then by tool B's scores, and look at whether the order matches before you look at how far apart any pair of numbers is. Two tools ranking a set differently is the disagreement that actually matters - two tools printing different absolute numbers on a matching order is usually just a scale offset.

What the procedure supports

If the rank order matches closely and the numbers differ by a roughly constant amount, that is a calibration difference between two tools that otherwise agree - useful to know, and not a reason to distrust either one. It is also why a single correlation figure is the wrong summary: Bland and Altman (1986), comparing clinical measurement methods, warned that "the use of correlation is misleading" when the question is whether two methods agree.

If the rank order itself shifts - a submission that placed third on tool A places seventh on tool B - that is a genuine disagreement about which submissions matter more, and it is worth digging into which axis is driving it.

If neither the order nor the gaps are stable across repeats of the same file, the noise floor of one or both tools is larger than the difference you are trying to measure, and no comparison built on a handful of submissions will be reliable until you know that floor.

Two mistakes that quietly break the procedure

Comparing a submission you already reused elsewhere. If the same file was previously used to test the first tool's settings, tweaked, cropped or filtered along the way, what reaches the second tool is no longer identical to what reached the first, even if you believe you undid the change. Keep a clean copy of each file aside specifically for the comparison and never touch it again.

Stopping after one pass. A rank order built from one run of five files can flip on a second run of the same five files if either tool has any noise in it at all, which most do. A short series repeated rather than a single pass is what tells you whether an order shift you saw was the tools disagreeing or one of them being noisy that day - running the procedure once and treating the result as final is the single most common way this goes wrong.

What this procedure does not do

It does not tell you which tool is right - there is no answer to that question in the sense most people want, only a set of narrower, checkable questions about consistency and disclosure that this procedure feeds into. It also assumes the submission itself was prepared the same way for both tools; standardising what you send in is a precondition for this procedure, not a step inside it. That precondition scales past two tools too - holding one fixed protocol constant while running it through several raters at once is the same discipline applied to a wider comparison.

The same logic extends past two AI tools. Comparing a model's result against a human reviewer's verdict is a different exercise - a person does not produce a repeatable number from an identical file the way a model does, so rank comparison across a small set works less cleanly there. And a comparison run against Rate Cock benefits from the same discipline as any other pairing - its own public entries are one place to pull a varied small set from, rather than relying on submissions of your own that may not span much of the scale.

Read next

Full archive