Tools

Why two tools disagree on the same photo

Submit one image to two raters and expect two answers. The gap is mostly calibration, and calibration is not accuracy.

3 min readTools

Run the same photo through two rating tools and you will get two different numbers, often two points apart. People treat this as evidence that one of them is wrong - as if two rulers had returned different lengths. But a ruler has a unit and a rating tool has a scale it invented, so the analogy fails at the first step. Usually neither is, in the sense they are being accused of.

Different scales wearing the same clothes

The dominant reason, and it is not close.

Both tools print a number out of ten. Underneath, each maps some internal value onto that scale by its own rule. One might map onto the distribution of everything it has ever scored - making its output a percentile. Another might map from a fixed range with the ends compressed.

Those two systems can rank a set of submissions identically and still print numbers two points apart on every one of them.

That is the test worth applying: not "do they agree on the number" but "do they agree on the ordering". Score three of your own photos on both tools. If both put them in the same order, the tools agree and only their scales differ. If the orders differ, you have found a real disagreement.

Different rubrics, different weights

A tool weighting surface condition heavily and a tool weighting proportion heavily will diverge on any submission that is strong in one and weak in the other - which is most of them, since the underlying properties do not correlate much.

This is a genuine difference of opinion rather than a scaling artefact - the closest an algorithm gets to two human judges disagreeing, which they do for richer reasons - and it is only visible if both tools show their components. With totals only, it is indistinguishable from the calibration case above. Whether a tool exposes its rubric is one of the few things that meaningfully separates them.

Different training distributions

Each scoring head learned from a particular set of images and human ratings. Different sets, different notions of what a high score looks like. That is the mechanism itself, described in more detail on AI Penis, and it is why "advanced AI" is not a differentiator: they all learned this way.

Nobody publishes theirs, so this is unfalsifiable from outside - worth knowing it exists rather than trying to reason about it.

Sampling noise

Real, and small. The same tool on a byte-identical input returns close to the same number. This is the least of it, and if two tools differ by two points, this explains approximately none of that.

Where the gap is widest

None of the four reasons above are constant across the scale. Two tools scoring the same set tend to diverge most at the top and bottom, and agree more in the crowded middle, which is worth knowing before reading a big gap on an unusually high or low result as a special kind of disagreement.

The disagreement that should worry you

Not a gap in the numbers. A gap in the direction.

If you improve the lighting and one tool goes up while the other goes down, at least one of them is not responding to what it claims to respond to. That is a real failure and it is detectable in about ten minutes with three photos and two tabs.

Whereas a consistent two-point offset, with both tools moving the same way when you change something, is two working instruments with different zero points.

How to get something useful out of two disagreeing tools

Pick one and stay with it.

Cross-tool numbers have no exchange rate, but within a single tool the comparisons are valid - the scale is at least consistent with itself. A within-tool comparison across three of your own photos tells you something. A cross-tool comparison of one photo tells you about the tools.

If you want to be able to check a tool's consistency at all, you need one that shows the components rather than a total, and ideally some public distribution to sit your number against. Rate Cock exposes the six-axis breakdown on public entries, which is enough to run that check without any special access - most tools publish nothing of the kind, and there is no way to audit a number you cannot decompose.

Read next

Full archive