Scores

Which number to keep from a series

People report their highest score; the median is the one that predicts what a fresh rating would say.

By 3 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

Keep the median, not the best: the middle of several ratings predicts what a fresh rating would say, while the highest mostly records which attempt the noise favoured. The more times you retry, the higher that maximum climbs, though the submission never changed.

Why the maximum is biased

Each rating is a draw from a distribution centred on some true value, with some spread around it either side, spread that comes partly from the model's own internal variance and partly from small differences between otherwise similar photos. Take one draw and it is an unbiased estimate of the centre. Take five draws and keep only the highest, and you have selected for the draw that happened to land furthest above the centre - the maximum of several draws is expected to sit above the true value, and the more draws you take, the further above it tends to sit. This is the same reason the best score you have ever seen from a tool looks better than the tool actually is - it was selected from many attempts, most of which nobody screenshots.

Why the median is not

The median does not care how extreme the most extreme draw was. It sits wherever the middle of the distribution actually is, which is close to the true value regardless of how many times you retried or how lucky one attempt happened to get. Take three ratings and the median is simply the middle one. Take five and it is the third when sorted. Either way, a single unusually high or low result pulls the mean but barely moves the median, which is exactly the property you want when one retake out of several might reflect a lighting slip or a bad crop rather than the submission itself. The NIST/SEMATECH Engineering Statistics Handbook puts it plainly: extreme values "distort the mean" but "do not distort the median since the median is based on ranks."

The practical rule

Take an odd number of ratings under matching conditions, sort them, keep the middle one, and treat the rest as evidence of spread rather than as candidates to discard. This does not tell you how many retakes are enough to trust the result - that is a separate question with its own answer. It also is not a method for deciding that a single wild result was a fluke worth dropping; that judgement call has its own procedure, and reaching for the median is not a substitute for it.

What it gives you is a number to report that does not quietly inflate every time you retry. The same reasoning applies past this one site: a headline result from any comparison tool, including one that reports several axes rather than a single figure, is more informative read as one draw from a series than as a verdict on its own. If the surprising number came from a measurement rather than a rating, the fix is different again - a length figure is only as good as the method that produced it, and retaking a rating tool does nothing for that. And if what you actually want is a considered opinion rather than a repeatable statistic, a human judge is not answering with a distribution at all, which is a different trade rather than a worse one.

Read next

Full archive