Photos
How to find out how much of your score is you
Identical file, retake, recrop, control shot, short series: the complete method for measuring how repeatable your results are and where the noise comes from.
Every score comes from a photo, and every photo carries some amount of noise that has nothing to do with the subject: a slightly different distance, a crop chosen by eye, a file that has been through a different chain of compression than the last one you sent. Most of that noise is invisible from a single submission, because a single submission has nothing to compare itself against. What follows is a set of five small tests, run from the submitter's side rather than the tool's, that together separate how much of your score's movement belongs to your process and how much belongs to something else entirely.
None of these tests requires special access to a tool, and none of them requires more than a handful of extra submissions. What they require is running them in something close to the right order, because each one isolates a narrower question than the one before it, and the narrower tests only make sense once the wider ones have ruled out the bigger sources of variance.
Test one: the identical file, twice
Start here, because it is the cleanest test available and it sets the floor for everything that follows. Submit the exact same image file to the tool twice, with nothing about it changed - not recompressed, not recropped, not resized. The reproducibility test covers the mechanics of this in full; the point for this method is what it establishes: whatever gap appears between the two results happened with zero difference in the input, so it cannot be attributed to anything about your photo. This is the tool's own noise floor, and every other test in this method is measured against it.
If this test returns a large gap on its own, that is worth knowing before doing anything else - a tool with a wide floor is going to make every subsequent test noisier too, and the honest response is patience with a larger sample rather than tighter photography.
Test two: a fresh retake, same conditions
Once you know the floor, reintroduce the one thing test one removed: take a new photo, trying to match distance, angle, light and crop as closely as you can to the last attempt, and submit that instead of the same file. The retake test is the full procedure. The gap this time is the floor from test one plus whatever slipped in your attempt to hold conditions steady - and because you already know the floor, the difference between the two gaps is a rough estimate of what your protocol itself is contributing. Separating these two sources properly is worth a longer look if the two numbers do not sit close together, because the gap between them is the part you can actually do something about.
Test three: three crops of the same photo
Take one existing photo and crop it three ways - tight, loose, and something between the two - then submit all three. Nothing about distance, angle, light or subject changed across the three submissions, so any spread in the results isolates framing sensitivity specifically, with none of the other variables mixed in. The recrop test walks through reading the spread; the short version is that a tight cluster means framing is a low-priority variable for your protocol, and a wide one means it deserves the same discipline as distance and angle already get.
This test costs nothing in new photography, which makes it worth running even when you are fairly confident about the answer - a confirmed low-sensitivity result is still useful, because it tells you where not to spend effort standardising.
Test four: the control shot
Set aside one photo, of the same subject under fixed conditions, that you never intend to update, and resubmit that same control shot alongside every new session going forward. The control shot covers the reasoning in full: because this image never changes, any movement in its score across sessions cannot be about the subject, and it becomes a running check on whether the tool itself has drifted since your last comparison.
This is the one test in the method that pays off over months rather than in a single sitting, and it is the closest thing available to a standing instrument calibration. A control shot that has drifted a full point over six months tells you something a single identical-file test run today cannot: not just that the tool has some noise, but that the noise has a direction, and that direction has been moving your comparisons for longer than you may have noticed.
Test five: a short series, summarised properly
For any single session where you actually care about the number - not testing methodology, just wanting to know where you stand today - take several photos under matched conditions rather than one, and summarise the set as a median and a range rather than trusting the first result. Five retakes summarised this way turns one draw from a noisy process into a centre with a stated spread, which is the only version of a rating result that is honest about its own uncertainty.
This test is where the previous four pay off, because by the time you run it you already know roughly how wide your protocol's noise is likely to be, from tests one through three, and you know whether the tool itself has stayed put since your baseline, from test four. A wide range on the short series is expected if tests one and two already showed a wide floor; a wide range when tests one and two were tight is a signal that something about this particular session went differently, worth checking against your log before drawing a conclusion.
Putting the five together
Run in order, these tests build outward from the narrowest possible question to the broadest. Test one asks whether the tool alone is stable. Test two adds your process into the same question. Test three isolates one specific variable within that process. Test four extends the whole question across time rather than a single sitting. Test five is where you actually use everything the first four established, to summarise a real result honestly instead of quoting a single number as if it carried no uncertainty.
A submission-side noise floor built this way is not a single figure - it is closer to a short profile: this tool's own instability is roughly this wide, my protocol adds roughly this much on top, framing matters or it does not, and the baseline has or has not drifted since last quarter. That profile is worth writing down once and referring back to, rather than rebuilding from scratch every time a result looks surprising.
What the profile is actually for
The point of building this profile is not to produce a number you show anyone else - it stays with you, as a reference for reading your own future results. Once you know roughly how wide the tool's floor is and how much your own protocol adds on top, a single surprising score stops being a mystery and starts being a checkable claim: is this difference bigger than the noise floor you already measured, or is it sitting comfortably inside it. Most surprising results, once checked this way, turn out to be ordinary variation dressed up as a finding, which is itself a useful thing to know before reacting to a number that moved.
The profile also ages. A noise floor measured today assumes the tool and your protocol both stay roughly where they were when you measured them, and neither is guaranteed to hold indefinitely. Revisiting tests one and four every few months, rather than treating the original profile as permanent, is what keeps the whole method honest over the stretch of time a real comparison actually needs to cover.
What this method does not cover
Everything here is about the submission and the tool's response to it, not about what happens inside the tool between input and output. What makes a hosted vision model's own inference vary on a technically identical input is a real and separate question, and it is the reason test one's floor is never exactly zero even for a tool with no bugs at all - some of that floor is architectural rather than procedural, and no amount of tightening your photography will close it. This method also assumes you have already fixed the conditions worth fixing in the first place; the actual protocol - which variables to record and how - belongs to a separate piece, and is worth reading before running these tests rather than after, since a protocol you have not written down cannot be checked against a suspicious outlier later.
Two adjacent categories run comparable exercises for different reasons. A physical measurement has its own repeat-and-average discipline built around a method rather than a model - taking more than one reading with a tape, and what to do when they disagree is a parallel practice, worth citing here because the underlying goal, a centre plus a stated spread rather than one number treated as exact, is identical even though nothing about a tape resembles a vision model. A human reviewer's session-to-session consistency is a third kind of question again, closer to attention and mood than to protocol or architecture - what shapes how a judge reads the same kind of submission across different sessions is worth a glance only to keep the three mechanisms - tool, tape, person - from collapsing into one imagined source of noise.
What five tests actually buy you
None of this makes a rating tool more accurate. What it buys is the ability to tell, the next time a number surprises you, whether the surprise is coming from your photo, from the tool's ordinary noise, from a drift the tool has picked up since your last comparison, or from something in your framing you had not previously flagged as a variable. That is a different and more useful outcome than a higher score, and it is available at the cost of a handful of extra submissions run once, kept as a reference, and checked back against periodically.
On a service like Rate Cock, where every submission stays logged against the account rather than shown once and discarded, this whole method is easier to run than it sounds - the five tests above are mostly a matter of choosing which existing or new submissions to compare, not building new infrastructure to track them. The result is a noise floor you actually know, rather than one you are guessing at every time a score moves and you are not sure why.