Tools
Telling a scale offset from noise
If tool B is always about a point above tool A, that is a scale difference; if it is sometimes above and sometimes below, that is something else.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Systematic disagreement means one tool sits consistently above the other; random disagreement means the gap changes sign and size from one submission to the next. Telling them apart takes only a handful of paired results laid out side by side.
The two kinds
The split is borrowed from metrology. The International Vocabulary of Metrology (JCGM 200:2012) defines systematic error as the component that, in replicate measurements, "remains constant or varies in a predictable manner", and random error as the component that "varies in an unpredictable manner".
Systematic. Tool B is consistently above tool A - a point, a point and a half, half a point, whatever the size - across every submission in the set. This is a calibration offset: the two scales are shifted relative to each other, but they still agree on which submission is stronger, and the offset itself can be roughly subtracted out once you know it is there.
Random. Tool B is above tool A on some submissions and below it on others, with no consistent direction or size. This is either noise - the floor of variation either tool would show against itself on a repeat - or a genuine divergence in what the two rubrics are responding to on different kinds of submission.
How to tell them apart
Take the handful of paired results from a proper two-tool comparison and look at the sign of the gap on each pair, not just its size. All positive, roughly the same magnitude: systematic. Some positive, some negative, no pattern: random.
A middle case is common and worth naming separately - a gap that is usually positive and occasionally crosses to negative by a small amount is probably a systematic offset with ordinary noise layered on top, and the reproducibility test run on either tool alone tells you how much noise to expect before you decide the crossing is meaningful.
Why the distinction matters
A systematic offset is close to harmless for most purposes. If you know tool B always scores about a point higher, you can read a result from either tool with that adjustment already applied mentally, and a comparison between two submissions on the same tool is unaffected by it either way.
Random disagreement is the one that should change how you use the tools. It means the two are not simply calibrated differently against the same underlying judgement - either one or both is noisy enough that a single result is not trustworthy on its own, or the rubrics are genuinely picking up on different things depending on the submission, which is closer to a per-axis divergence than a scale issue. A physical measurement has no equivalent of this distinction at all: a length taken by a documented method either matches a second measurement within instrument error or it does not, there is no rubric underneath it to diverge.
A worked pair
Five submissions, tool A and tool B: 6.0/7.1, 6.4/7.5, 7.0/8.0, 7.3/8.3, 7.9/8.9. Every gap sits between 1.0 and 1.1 - systematic, and close enough to a flat offset that subtracting roughly one point from tool B brings the two into rough agreement on every submission. Now take the same five with gaps of +1.1, -0.4, +0.8, -0.6, +0.3. No consistent sign, no consistent size - that is random, and averaging the gap here would produce a number that does not describe any single pair, only a meaningless mean over a scatter. The first pattern is worth a mental adjustment; the second is worth treating both tools' single results with more caution than either would otherwise deserve.
Where this sits next to the rest
This question comes after you already have paired results in hand, not instead of collecting them - the procedure for generating the pairs is separate from reading them, and this piece only covers the reading. It is also a narrower question than deciding which tool is right, since telling systematic from random disagreement does not require settling correctness at all.
Systematic and random are useful categories outside a rating tool comparison too. A single human review has no second result to compare against, so this distinction does not translate directly - a reviewer's one verdict cannot be sorted into systematic or random against itself. Where the underlying disagreement traces back to how two different models process the same image, that is a property of the systems themselves rather than of the rubric sitting on top of them, and it is worth knowing which one you are actually looking at. Rate Cock reports six axes rather than a total, which gives a systematic-versus-random check somewhere finer than the total to run on - an offset that holds on five axes and reverses on one is a different finding from an offset that holds everywhere.