Photos
What a series buys and what it does not
People submit again hoping for a higher number; what more submissions actually deliver is a tighter estimate of the same number.
Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you
More photos do not raise the score; they sample the tool's noise more times. Retry often enough and one attempt will land higher, but the highest of several is just the top of the range, not a truer number. Treating it as the real score is the most common way people misread their own submissions.
The misconception
More tries, better score - as if each submission were an independent shot at improvement, and the best of them represents what the tool "really" thinks. That model would make sense if each attempt were meaningfully different, but under a standardised protocol the underlying subject and conditions are the same each time. What varies between attempts is mostly the tool's own noise: minor shifts in how it reads a nearly identical image, plus small photographic differences a person cannot fully eliminate by hand.
What a series actually delivers
A tighter estimate of one underlying value, not a series of independent judgements trending upward. Five submissions under the same fixed conditions give a centre and a spread around that centre, and the centre is the useful number - it is closer to what the tool would say on average than any single result, including the highest one.
This only works if the conditions are actually held constant. A set of photos where distance, light and angle drift between shots is not a series in this sense; it is several different, uncontrolled submissions, and averaging or maximising across them tells you about the drift as much as about the tool. How rating tools differ in what they even show you - a total only, or a breakdown - also affects how much a series can teach you, since a breakdown lets you see whether a given axis is moving with the noise or moving for a reason.
Why the maximum is the wrong summary
Selecting the best of several tries and reporting only that number is a form of selection bias: it systematically overstates what the process produces on average, and it gets worse the more attempts are made, because a larger sample is more likely to contain an extreme value somewhere in it by chance alone. A maximum reported without the attempts behind it looks like a single clean result and is actually the peak of noise. Epidemiologists call what happens next regression to the mean: Barnett, van der Pols and Dobson (2005) describe how "unusually large or small measurements tend to be followed by measurements that are closer to the mean", and note that the effect grows with measurement error.
The same reasoning applies whether the number came from a model or a person scoring the submission - a favourite result pulled from several attempts is not more representative just because a human produced it, and what a repeated measurement is supposed to converge on is the same statistical idea in a physical context, where nobody would report only their tallest single reading either.
Why more attempts make the problem worse, not better
It is tempting to think that submitting many times is at least harmless, even if the maximum is the wrong summary to report. It is not harmless if the maximum is what gets remembered or shared, because the gap between the maximum and the true centre grows with the number of attempts. Ten tries pull the top value further from the centre than three tries do, simply because a larger sample has more room to throw up an extreme. The more someone resubmits chasing a higher number, the less representative the number they eventually land on and keep becomes.
What to do instead
Fix the conditions, take several submissions, and use the centre of the results rather than the best one. That centre is the number worth writing down, the number worth comparing to a future series, and the honest answer to what the model's actual read of the submission has been converging toward the whole time. A service like Rate Cock, which keeps a history rather than showing only the latest result, makes this easier to see for yourself: what a stored score is meant to represent is spelled out there directly. It will usually be lower than your best attempt. That is not the tool being unkind; it is the best attempt being an outlier, which is exactly what a highest-of-several value is built to be.