Tools
When the totals agree and one component does not
Disagreement on a single axis is more informative than disagreement on the total; it points at a rubric difference you can name.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Per-axis disagreement - two breakdowns that match everywhere except one axis - is more informative than two totals that differ, because it narrows the cause to how that single axis is defined or read. Two disagreeing totals could mean almost anything.
Why isolation matters
A total is a sum of several judgements, so when two totals differ you cannot tell whether one axis moved a lot, several moved a little, or the weighting rule itself is different between the two tools. The total alone cannot answer which of those it was.
A single divergent axis inside two otherwise-matching breakdowns rules most of that out at once. If four axes line up closely and a fifth is a point and a half apart, the weighting is probably similar, the general judgement is probably similar, and the difference sits in how that one axis specifically is being defined or read.
How to find it
Run the same submission through both tools and lay the breakdowns next to each other, axis by axis rather than total against total. Axis names that look identical across tools are not guaranteed to mean the same thing, so line them up by what each axis actually claims to measure, not just by the label printed next to it.
Do this with more than one submission before drawing a conclusion. A single comparison could catch either tool on a noisy day on that one axis; a divergence that shows up the same way across several submissions is the pattern worth naming - how the model behind either tool actually processes an image is a separate question from why the two rubrics diverge, and it is easy to blame the wrong one.
With enough paired submissions, agreement on each axis can be put on a formal footing. Clinical reliability research uses the intraclass correlation coefficient, and Koo and Li (2016) grade it as poor below 0.5, moderate from 0.5 to 0.75, good from 0.75 to 0.9 and excellent above 0.9 - while warning that the 95% confidence interval, not the point estimate, should decide the grade.
What to conclude from it
A consistent single-axis gap usually means one of two things: the two tools weight the same underlying signal differently within that axis, or they are naming two genuinely different things the same label. Either way, the finding is specific enough to act on - you can discount that one axis when comparing the tools, or read it with the knowledge that it is measuring something narrower or broader than its name suggests on the other tool.
What it does not mean is that either axis is wrong. Neither tool has ground truth to be wrong against; the disagreement tells you the two rubrics differ at that one point, not which rubric is correct.
A worked shape
Say two tools each report proportion, symmetry, presentation and texture, and both land within a few tenths of each other on three of the four. On the fourth, one prints a 5 and the other an 8. That is not a rounding wobble - it is large enough to be the finding itself, and the next step is not to average the two but to read what each tool's rubric actually claims that axis covers. Sometimes the label hides two different questions: one tool's "presentation" may fold in grooming and framing together, while the other keeps grooming under a separate axis entirely, and a single divergent number is often just that split showing up as a gap.
Where this connects
This is a narrower question than why two tools disagree on the same submission at all, which covers calibration and scale differences that show up in the total as well as any single axis. It is also distinct from mapping one tool's full set of axes onto another's - that is an attempt to build a general translation, where isolating one disagreement is a specific, narrower check you can run in minutes.
The same isolating logic works outside rating tools too: if a human reviewer and a model agree on the overall impression but diverge on one specific point, that divergence names something worth asking about directly rather than dismissing as a difference in overall taste. And a rubric axis is not the same kind of measurement as a length taken with a tape - the axis is a judgement, not a physical quantity, which is exactly why two tools can disagree on it with neither one being wrong. Rate Cock reports six axes per submission, which is enough separate components that a per-axis comparison against another multi-axis tool has somewhere real to land.