Tools
The tools that compare instead of score
A pairwise tool never claims a value, only an order; that is a weaker claim and a more defensible one.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Most rating tools answer "how good is this." A pairwise tool answers a smaller question: given two submissions, which one wins. That is a genuinely different output shape from the other four kinds of rating tool, and the difference is worth taking seriously rather than treating a pairwise result as a score with extra steps.
What it outputs
Two submissions go in, and one verdict comes out: A over B, or B over A. There is no number attached, no scale, no claim about how much better - just an order between exactly the two things compared. Some pairwise tools run this repeatedly against a fixed reference set or against other users' submissions, building up something that looks rank-like over many comparisons, but each individual judgement the tool makes is still binary: which one, not how much. The approach is established well beyond this category: the Chatbot Arena project (Chiang and colleagues, 2024) ranks large language models from crowdsourced pairwise votes, over 240,000 of them at the time of the paper.
Why order is a more defensible claim than value
A scored result claims a position on an entire scale - it says "this is a 7.2," which is a strong claim requiring the tool to have calibrated a whole range correctly, from bottom to top, and to place this particular submission accurately within it. A pairwise result only claims that, between these two specific things, one edged out the other. That is measurably less for the tool to get wrong. It does not need a stable zero, a stable ceiling, or a consistent mapping across its entire range - it only needs to be reasonably consistent on the one comparison directly in front of it, which is a smaller, more tractable problem than scoring absolutely.
This is the same reason ranking a set of your own submissions is something a tool does more reliably than scoring any one of them - relative judgements throw away less honesty than absolute ones, because they never have to commit to what a 10 would mean, only to which of two things is closer to it.
What it cannot tell you
The obvious cost of the smaller claim is that a pairwise tool cannot answer "how good," only "better than what." Win a hundred pairwise comparisons against weak reference images and the result says nothing about an absolute standard - it says you beat what you were compared against, and the value of that statement depends entirely on what the reference set was, which most pairwise tools do not disclose in detail. There is also no way to read a "close win" as meaningfully different from a "clear win" unless the tool separately reports a margin, and many do not, collapsing what may have been a near-tie into the same binary verdict as a landslide.
Order itself can leak into the verdict when a model is the judge. Wang and colleagues (2023) found that simply swapping which answer appeared first let Vicuna-13B beat ChatGPT on 66 of 80 queries with ChatGPT as the evaluator, and recommended aggregating verdicts across both orderings. A pairwise tool that does not say it runs each pair both ways has left that bias unaddressed.
How this differs from rank-only tools
Pairwise and rank-only can sound similar - neither hands back a scale number - but they are not the same claim. A rank-only tool positions one submission within a whole population at once, an aggregate statement built from many comparisons the tool already ran on your behalf. A pairwise tool makes exactly one comparison at a time, visible and specific, and any aggregate position has to be built up deliberately by running many of them, which most users never do - most encounters with a pairwise tool are a single verdict on a single pair, nothing more.
Where it is useful
Pairwise comparison is well suited to exactly the situation a score is bad at: choosing between two of your own submissions when you cannot tell by eye which one is stronger. Because the tool only has to be locally consistent rather than globally calibrated, a pairwise verdict between two close candidates is often more trustworthy than the half-point gap between their two absolute scores would be on a single-number tool.
What sits outside this type
None of this maps onto what an underlying model can actually resolve at the pixel level or onto a measured comparison between two states - both are inputs a pairwise tool might use, not versions of the pairwise claim itself. It also has a human equivalent worth naming rather than covering here - asking a person to pick between two submissions is the same output shape, order rather than value, arrived at by a different process entirely. A pairwise verdict from either source answers the same narrow question, and neither one should be read as a stand-in for the kind of number Rate Cock reports on a full breakdown - the two outputs are structurally different claims, not two ways of expressing the same result.