Tools
Same file, twice, and what the gap tells you
Submit the identical file twice; the difference between the two numbers is the floor of the tool's noise, and it is the first thing worth knowing.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Before judging whether a rating tool's number is right, check whether it is even the same number twice. Take one file, submit it, note the result, submit the exact same file again as a fresh upload, and compare. That is the entire test, and it takes about as long as reading this sentence twice.
What this test isolates
Nothing about the submission changes between the two runs - not the distance, not the light, not the crop, not the file itself. Any difference between the two results is therefore coming from inside the tool: its own internal variance, a non-deterministic model call, or a calibration step that is not perfectly stable run to run.
This is a narrower question than whether two different photos of the same subject score the same. It is narrower on purpose - by holding the file itself identical, the test removes every submission-side variable at once and leaves only the tool's own consistency to measure. Metrology would call this a repeatability check: the International Vocabulary of Metrology (JCGM 200:2012) defines repeatability conditions as the same procedure, measuring system and operating conditions, with replicate measurements on the same object over a short period of time.
Running it properly
Use the exact same file, not a re-export or a re-save - re-encoding a JPEG, even at the same quality setting, can change the bytes the tool receives and confound the result. Submit it as two separate uploads rather than refreshing a cached result page, since a refresh may just be showing you the same stored answer rather than a fresh evaluation.
Space the two submissions out a little if the tool allows it - back to back within the same session risks catching a cache rather than a genuine second evaluation. A few minutes apart is enough for most tools; if a service documents a caching window, wait past it.
Reading the gap
Zero, or effectively zero. The tool is either deterministic on identical input or caching the result rather than recomputing it. Either is a defensible design choice, but they are different claims - a deterministic model genuinely reproduces, a cache only reproduces its own memory of the first answer, and telling them apart usually needs the tool's own documentation or a support answer, since both look the same from outside.
A small, consistent gap - a few tenths. Normal. Most hosted vision models carry some run-to-run variance, and a small gap here is the honest floor of noise that every other result from this tool inherits. This is the number worth writing down: it is your yardstick for whether a later change in your own score is real or inside the tool's own wobble.
A wide gap - most of a point or more. A finding, not noise. A tool that cannot return a consistent number on an unchanged file has a floor of noise wide enough to swallow most real differences you might be trying to detect, and any comparison you run against it - before and after, this tool versus that one - inherits that width whether you account for it or not.
What this test does not measure
It says nothing about accuracy - a tool can be perfectly consistent and simply wrong, or scoring something other than what it claims to. Consistency and correctness are separate properties, and this test only speaks to the first one. It also has no equivalent on the human side: a person reviewing a submission is not going to give the identical response to the identical file twice, and that is not a flaw to test for - a human reviewer is a different kind of output, not a machine that failed a consistency check.
It also is not the retake test. Submitting a fresh photo taken under the same conditions measures your submission noise plus the tool's noise together - useful for a different question, about how stable your own results are, but it will not isolate the tool the way an identical file does. The words "repeatable" and "reproducible" get used interchangeably here but describe two distinct claims, and knowing which one a given test is actually making avoids overreading a passed check. Run the identical-file test first to establish the tool's floor, then the retake test to see how much extra variance your own conditions add on top of it.
Where the variance actually comes from
Some of what shows up as tool-side noise on this test is inherited from the model layer itself rather than anything the results page controls - the same hosted vision models that sit underneath most rating tools carry their own run-to-run variance, independent of any calibration step layered on top. A tool cannot fully eliminate that floor; the honest ones acknowledge it rather than presenting every result as a fixed, exact reading. A tool that reports per-axis components rather than one blended figure makes this test more informative rather than less - run it against a breakdown like Rate Cock publishes on public entries and you can see which specific axis carries most of the run-to-run wobble, instead of only a single total moving for reasons you cannot localise.
Using the result
Once you know the floor, use it as a lower bound on what counts as a real difference elsewhere. A change of a few tenths between two of your own submissions, on a tool whose identical-file gap is itself a few tenths, is not yet distinguishable from that floor - how big a gap has to be before it is worth treating as real depends directly on the number this test gives you.
This is also the fastest single check to run before trusting a tool you have not used before, and it costs one extra upload. The rest of the evaluation covers the rubric and the disclosure around it, but consistency is the one property you can confirm yourself, on the spot, without reading a single page of documentation - and a tool whose repeat number is embarrassing tends not to advertise that this test exists. For a submission-side length claim, the equivalent discipline lives with the tape rather than the lens - a measurement records its own repeat method separately, and that is a different kind of reproducibility test entirely.