Photos
Separating two sources of variance
Two tests separate them: the same file twice measures the tool; a fresh photo under the same conditions measures you plus the tool. Subtract.
Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you
To tell tool noise from photo noise, run two tests: submit the identical file twice, which isolates the tool, then a fresh photo under the same conditions, which measures you plus the tool. The difference between the two gaps is roughly what your setup contributes. A single comparison mixes both sources and cannot separate them.
Two tests, two sources
The first test removes your variation entirely: submit the identical file twice. The reproducibility test covers how to run it and what counts as a pass, but the short version is that whatever gap appears between the two results is the tool's own noise floor - nothing about your photo changed, so nothing about the photo can explain it.
The second test reintroduces your variation on purpose: take a fresh photo under conditions you tried to hold constant, and submit that. The retake test walks through the setup in detail. The gap this time is tool noise plus whatever slipped in your protocol - a slightly different angle, a half-step of distance, light that shifted five minutes later.
Run in isolation, either test tells you something. Run together, they tell you where a discrepancy actually lives.
Reading the two gaps against each other
Call the identical-file gap the tool's floor and the retake gap the combined figure. Measurement science splits variance the same way: the NIST/SEMATECH Engineering Statistics Handbook treats repeatability, estimated from repeated measurements of one item, as "the basic precision for the gauge," separate from variation added between occasions. The retake gap can never be smaller than the floor - you cannot remove noise by adding a variable - so if the two numbers come back close together, your protocol is not the problem. Whatever variance is showing up in your results belongs to the tool, and no amount of standing more carefully will shrink it.
If the retake gap is meaningfully larger than the floor, the difference between the two is a rough estimate of what your setup is contributing on its own. That number is worth chasing down, because it usually points at one of the variables standardising a submission is meant to pin down - distance drifting a few centimetres, the angle not quite level, the crop landing differently by eye each time. Sometimes the culprit is upstream of the framing entirely - how exposure and ISO settings introduce grain a tool then has to interpret is worth ruling out before blaming the protocol for a gap the camera itself produced.
A retake gap close to zero while the identical-file gap is not near zero is the case people find confusing, and it is also informative: it means the tool responds to something in re-encoding or re-uploading a file that has nothing to do with the image content, which is a property of the tool's pipeline rather than a rubric decision.
What this does and does not fix
Neither test tells you whether the score is right. Reliability research draws the same line; Koo and Li (2016) grade test-retest agreement from "poor" below 0.5 to "excellent" above 0.9 on the intraclass correlation, and none of those grades says anything about accuracy. Both only tell you how much of the number's movement to attribute to which side of the upload. That distinction matters because the fix is different in each direction - a noisy tool is addressed by running a short series and reading the median rather than trusting one draw, while a noisy protocol is addressed by tightening the conditions themselves.
The same logic applies outside this category. Whether the tool's underlying model is stable in the first place is a question about what happens before any score reaches the display, and it is the reason the identical-file test can return a nonzero gap even when nothing about the process looks wrong from the outside. A physical measurement does not have this ambiguity built in the same way - a tape read against a fixed method either repeats or it does not, and there is no separate "tool noise" term to isolate, because the instrument is not doing any interpretation. A human reviewer's variance is a third kind again, closer to mood and attention than to either of these - what shapes a judge's read from one session to the next is a different mechanism, worth naming so it is not confused with either test here.
Keep both gaps in mind before you draw a conclusion from any single result. On Rate Cock, where results are logged per entry rather than shown once and discarded, it is possible to run this same identical-file and retake comparison against your own history instead of guessing from memory - the size of gap worth taking seriously is smaller than most people assume, and knowing which side of the upload it came from is most of the work of interpreting it.