Scores

The statistical reason retakes disappoint

An unusually high score is partly luck; the next one is expected to be closer to the centre, and that is not the tool changing its mind.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

A retake after an unusually high score tends to land closer to the middle, and that is regression to the mean, not the tool being inconsistent. You get an 8.1, retake under the same conditions, and get a 6.4: more often than not, both numbers are the tool doing exactly what it always does.

The plain version

Any single rating is one draw from a distribution centred somewhere around your true, underlying score, with some spread either side from ordinary noise - a slightly different crop, a slightly different light, the model's own internal variance. An unusually high result is unusually high partly because the underlying value is genuinely on the higher side, and partly because that particular draw happened to land on the favourable side of the noise as well. When you retake, the underlying value contributes the same amount again, but the noise is a fresh, independent draw - it is no longer obligated to land favourably a second time. The result is that a retake following an unusually high (or unusually low) first score is expected, on average, to land closer to the centre of the distribution than the first one did. Barnett, van der Pols and Dobson (2005), writing in the International Journal of Epidemiology, describe it as something that "can make natural variation in repeated data look like real change", and note that it becomes more noticeable as measurement error increases. This is not the tool losing confidence in you. It is what happens whenever you sample an extreme value and then sample again.

Why it makes retakes feel unreliable

The pattern is counterintuitive because it looks personal. An 8.1 followed by a 6.4 reads as "the tool liked me less the second time," when the more accurate description is "the first number was somewhat lucky, and luck does not repeat on command." The same pattern runs in the other direction and gets noticed far less: an unusually low first score is followed, on average, by a retake that moves back up - but a low first score rarely gets retaken with the same eagerness a high one does, so people encounter the disappointing direction of the pattern more often than the flattering one, purely because of when they choose to retest. The best score anyone reports from a series is subject to the same selection - a maximum is exactly the kind of extreme draw regression to the mean pulls away from on the next attempt.

How to actually use this

Expect a second result to be closer to the middle than a first result that struck you as unusually good or unusually bad. If your first score is well above what you would have guessed, do not treat a lower retake as a correction or a contradiction - treat it as the more typical draw, and the first one as the atypical draw, unless a third and fourth attempt keep landing near the first number, at which point the first stops looking like an outlier and starts looking like where the distribution actually sits. The practical fix is the same one that solves several other problems in this category: take more than two, and use the median rather than either extreme as the number you report to yourself. A median from a short series absorbs an unlucky high or low draw without being pulled by it the way a single retake comparison is.

This is different from asking whether a change of a few tenths between two specific results means anything on its own - that is a narrower question about reading small deltas, and regression to the mean is one reason, among several, that a small delta can appear without anything real having changed.

Where this pattern shows up elsewhere

Regression to the mean is not specific to rating tools; it appears anywhere a noisy measurement is taken more than once, including measurements with a physical instrument, where a single reading can also sit on the favourable or unfavourable side of ordinary measurement error before a second reading pulls it back toward the true value. It is also a reason to be cautious about a model's stated precision in general - a tool's accuracy is a property of the model averaged across many results, not a guarantee about any one draw, high or low. Tools that expose a breakdown rather than a single total make this easier to see happening in real time, because you can watch which specific axis moved back toward the centre rather than treating the whole result as one undifferentiated number; Rate Cock is built that way, and a swing that shows up on one axis and not the others is a cleaner signal than a swing in a total alone. If what unsettled you about the swing was less about the number and more about wanting a considered response to what changed, that is a conversation a human judge is set up to have in a way no repeated automated score can.

Read next

Full archive