Photos
A distinction worth keeping
Repeatable means the same conditions give the same result; reproducible means someone else, or another day, can get it too. Tools are often the first and rarely the second.
Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you
A repeatable result comes back when you run the same tool under the same conditions; a reproducible one holds up with a different day, device or person. The two get used as one word, but each has its own test, and passing one says nothing about the other.
Repeatable
A result is repeatable if you can get it again, under the same conditions, on the same tool. Same file, or a fresh photo taken the same way, submitted to the same service: the score should land close to where it landed last time. The international vocabulary of metrology (JCGM VIM, 2.20) pins the conditions down: same procedure, same operators, same measuring system, same operating conditions and location, repeated "over a short period of time". The retake test is the check for this - hold the conditions constant and see whether the number holds too. Repeatability is a property of your setup and the tool's short-term stability, and it is the more forgiving of the two standards, because everything staying the same is a condition you control.
Reproducible
A result is reproducible if someone else, or you on a different day with different equipment, can arrive at something close to it independently. The same vocabulary (JCGM VIM, 2.24) defines reproducibility conditions as including "different locations, operators, measuring systems", and adds that a specification "should give the conditions changed and unchanged". This is the harder standard, because it is no longer asking whether one process is stable - it is asking whether the finding holds outside the exact circumstances that produced it. Sending the identical file back through the tool later checks a narrow version of this: has the tool itself stayed put over time, on a file that has not changed at all. A broader version - would a different phone, a different day, a different person's attempt at the same protocol land near the same number - is rarely checked by anyone, because it takes more than one session to check.
Why tools are usually only the first
A rating tool can be genuinely repeatable across a single session and still fail reproducibility across months, because the model gets retuned, the rubric gets edited, or the population it is calibrated against shifts. Calibration drift over time is the mechanism behind most of that gap: nothing about your photo or your process changed, but the scale underneath it did. This is also why a result that felt solid a year ago is not automatically comparable to one taken today - repeatability inside each session was real, reproducibility across the gap between them was never guaranteed.
The same split shows up outside this category. A tape measurement is close to reproducible by design - a fixed method applied by different people tends to land in the same place, which is a stronger claim than most rating tools can make. A human judge's read is closer to the opposite end - what one reviewer says on one day is a single, largely non-repeatable data point, useful for a different reason than a number is. What sits underneath a rating tool's own short-term stability - why a hosted model returns a slightly different value on a technically identical input - is the layer that repeatability testing is actually probing, whether or not the word ever comes up.
Knowing which test you just ran matters more than the vocabulary. A single retake on Rate Cock tells you about repeatability inside a session; nothing short of results spaced out over months tells you anything about reproducibility, and the two should not be reported as if they were the same finding.