Photos
The case for a short series every session
A single submission is one draw; five under identical conditions give you a centre and a spread, and the spread is the part that keeps you honest.
One score is a single draw from a process that has noise in it, and treating a single draw as the answer throws away the one piece of information that would tell you how much to trust it. Five draws, taken under the same conditions in the same sitting, give you two things a single submission cannot: a centre that is less affected by any one unlucky result, and a spread that tells you how wide the uncertainty around that centre actually is. Neither number is available from one shot, no matter how carefully that shot was taken.
Why five specifically
Fewer than three does not show you a spread worth trusting - two numbers just show you a gap, not a shape, and you cannot tell an outlier from ordinary variation with only two points. More than about five or six runs into diminishing returns for the amount of retaking involved: the centre stabilises quickly and each additional shot buys less certainty than the one before it. Five sits at the point where you can see a shape - most results clustered, maybe one that sits apart - without turning a rating session into a photography session.
The conditions have to genuinely hold across all five, which is the part people skip. Fixing distance, angle, light and crop before the series starts is what makes the five results comparable to each other in the first place; five photos taken under five different conditions are not a series, they are five unrelated data points that happen to share a session.
Summarising the series
Report the median, not the mean. A mean lets one extreme result pull the whole figure toward it, and rating tools do produce the occasional extreme result for reasons that have nothing to do with the subject - a processing hiccup, a marginal read on one axis. The median sits at the middle of the ordered set and barely moves when one value is unusual, which is exactly the property you want from a summary number.
Report the range alongside it, not instead of it. The median without a range looks exactly as precise as a single score, which defeats the purpose of running five in the first place - the range is the evidence that the median is a centre and not a coincidence. A tight range across five retakes is itself informative: it tells you the tool and your protocol are both behaving, on this particular session, and that the number is worth quoting elsewhere.
What a series with error bars gets you
A single number invites a single kind of comparison - is it higher or lower than last time - and that comparison is unreliable exactly because a single number does not carry its own uncertainty. A median with a range invites the honest version of that comparison: whether two ranges overlap is a better question than whether two points differ, because it accounts for the fact that neither point was ever going to be exact. Turning one result into a range this way is the same move used to compare a result against an earlier session, months apart, where the conditions were never going to be identical anyway.
This is a procedural habit, not a mathematical one, and it costs five submissions instead of one. It is also the fastest way to notice when something upstream has changed - the noise itself can live in the photo or in the tool, and a series taken this session against a series taken last session is how that distinction becomes visible rather than theoretical.
The same habit exists elsewhere in the category, in different form. A measurement taken with a tape gets its own kind of repeat-and-average discipline - the method for taking more than one reading is a parallel practice with a parallel goal, spread and centre rather than one number treated as exact. A model's own internal consistency across repeated inputs is a related but separate question - what makes a hosted vision model return a stable answer twice sits underneath every series you run, whether or not you ever think about it. And a human reviewer, asked the same thing twice, brings a version of this same variance for entirely different reasons - how a judge's attention holds across a session is worth naming here only because the underlying question, does one read tell you enough, is the same question in a different register.
Five retakes cost more than one. What they buy is a number you can defend, not just a number you can share - and on a service like Rate Cock, where every entry is logged rather than shown once and forgotten, a session's worth of results stays available to compare against the next one exactly this way.