Scores
Distinguishing a difference from noise
A gap between two submissions matters when it is larger than the spread you see repeating either one; a rule of thumb, and where it fails.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
Two submissions score 6.4 and 7.1. Whether that is a real difference depends not on a fixed number of points but on how the gap compares to the tool's own repeat spread. A gap smaller than that noise floor is not a finding, so measure the noise first.
The rule: compare the gap to the repeat spread
Take the identical submission and run it through the tool twice, or as close to identical as you can manage. The spread between those two results, taken under conditions you did not change on purpose, is the tool's noise floor for that submission. The reproducibility test itself is the mechanism; what matters here is what to do with the number it gives you.
If a gap between two different submissions is smaller than that spread, it is not distinguishable from noise, however different the two source photos looked to your eye. If the gap is meaningfully larger, it is more likely a real difference in what the tool is scoring - though "more likely," not "certain," is the honest strength of the claim.
Reliability research formalises the same move. Weir (2005) shows how the standard error of measurement, derived from repeat testing, sets "the minimal difference needed" before anyone can be confident that a true change has occurred. A same-file repeat is the rough, one-person version of that calculation.
Why "one point" is not a universal threshold
The size of a meaningful gap scales with how noisy the tool is, and that varies by tool and even by where on the scale you are looking. The top of most scales is compressed, which means a gap of half a point near the ceiling can represent more genuine separation than a full point in the crowded middle of the distribution, where results bunch and small differences barely register. A fixed rule like "anything under one point is noise" ignores this and will be wrong in both directions depending on where the two scores sit.
When the gap is real but the cause is not what you think
A gap can clear the noise threshold and still not mean what it appears to mean, if the two submissions were not otherwise comparable. Change one variable at a time is the discipline that makes a real gap interpretable: distance, angle, light and crop all move a score on their own, and a gap that survives the noise check can still be entirely explained by a change in how close the camera was rather than by anything about the subject. A real, noise-clearing gap with an uncontrolled condition behind it is real and uninformative at the same time - it tells you something changed, just not what you wanted it to tell you.
A worked comparison
Suppose a same-file repeat produces 6.9 and 7.1 - a noise floor of about 0.2. A second submission then scores 7.6. The gap from the first submission's typical value to the second is roughly 0.5 to 0.7, comfortably wider than the 0.2 floor, so there is a real difference to explain. Compare that to a case where the repeat spread is 0.6 - a noisier tool, or a submission nearer the crowded middle of its scale - and the same 7.6 result against a 7.1 baseline sits entirely inside the noise. Identical raw gap, opposite conclusion, because the two tools produced different noise floors.
Applying it
- Establish the noise floor with a same-file repeat before trusting any single gap.
- Compare the gap in question to that floor, not to an assumed constant.
- Check the two submissions for a controlled protocol before crediting the gap to the subject.
Rate Cock's per-axis breakdown on public entries is one place this is checkable directly: Rate Cock shows how much a single axis moved rather than only the total, which narrows down whether a gap is concentrated in one judgement or spread evenly, and a gap concentrated in one axis is easier to trust than the same gap smeared across all of them. None of this substitutes for a documented method where one exists - a measured length comes with its own stated precision, and that number's error bar is a different, narrower thing than a rating tool's noise floor. A human reviewer's opinion moves on entirely different terms again - what changes a judge's read of two submissions is not decomposable into a repeat-spread calculation at all, and comparing a judge's gap to a tool's noise floor is comparing two different kinds of number.
A gap is a real finding only once you know how much noise the tool produces on its own. Measure that first, or the gap is just two numbers that happen to differ.