Scores
Using scores to order your own submissions
A tool orders a set of your submissions more reliably than it scores any one of them; the ends of the order are trustworthy, the middle is not.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
Ranking a set of photos by score is reliable at the ends and unreliable in the middle: the clear best and clear worst usually separate beyond the tool's noise, while close results do not. A single score is a noisy estimate of one submission, and several scores from the same tool produce something sturdier - an order.
Why relative order beats absolute score
An absolute score carries the tool's full calibration burden: where the scale is centred, how compressed the top is, what the mapping assumes about the population. An order between your own submissions, scored by the same tool under the same conditions, cancels most of that out. Whatever the calibration is, it is applied consistently across the set, so a higher output for one submission relative to another is evidence about the relative comparison even when the absolute numbers are not trustworthy on their own. This is the same logic behind why a within-person comparison holds up better than a between-person one - fixing more of the variables that are not the thing you are measuring makes the comparison that remains more informative.
Where the ordering is reliable
The ends of a ranked set are the most trustworthy part of it. The submission that scores clearly highest and the one that scores clearly lowest are usually separated by a gap that exceeds the tool's noise floor, and a gap that clears the repeat-spread test is a gap worth taking seriously.
Where it breaks down
The middle of the set is a different story. Submissions that land close together in score are often inside the tool's noise band relative to each other, and their relative order can flip on a retake without anything about the submissions changing. Treating adjacent middle-of-the-pack results as a meaningful ranking - "this one beat that one by a tenth of a point" - reads precision into a comparison the tool cannot actually support at that resolution. Statisticians found the same thing with real league tables: in Marshall and Spiegelhalter's 1998 BMJ analysis of 52 UK fertility clinics, there was "great uncertainty" about true ranks even where clinics' rates differed significantly, and many clinics changed rank between years although their rates did not change significantly. A human reviewer ranking the same set works from a different kind of judgement entirely - a judge comparing several submissions is not bound to a noise floor the way a rating tool's output is, which is a different strength and a different limitation, not a substitute for either.
A practical read on a five-item set
Take five submissions scored 6.1, 6.9, 7.0, 7.1, 7.8. The top result and the bottom result are separated from their nearest neighbours by enough to trust: 6.1 is clearly last, 7.8 is clearly first. The three in the middle - 6.9, 7.0, 7.1 - span a range narrower than most tools' repeat noise, and re-running any one of them could easily shuffle that order. Reading them as third, fourth and fifth in a confident sequence is the exact overreach this technique invites; reading them as a tied middle is the accurate version of the same five numbers.
Using it well
Rank a set for the purpose it is good for: identifying your clear best and clear worst results out of several, not litigating the order of everything in between. If two results are close, treat them as tied rather than trusting whichever printed a marginally higher number. Building a set that is actually comparable in the first place - same conditions across every entry - is a precondition for any of this; an order across submissions taken under different setups is ordering the setups as much as the subject.
Rate Cock's own guidance on reading an individual result explains what your score means for one submission at a time; ranking a set is the tool for a different question - not what one number says, but which of several attempts actually separated from the rest. The same distance-and-angle discipline that keeps a set comparable is the discipline aipenis.com covers for accuracy generally - a set ranked against a drifting protocol is not really a ranking of the subject at all. None of this is available on the physical side, where a tape gives you a single documented figure rather than a ranked set to work with - ranking is a rating-tool technique because the noise it is designed to average out is specific to how these tools produce a number.
A ranked set tells you your best and your worst with more confidence than any single score gives you either one. It does not tell you much about the middle, and treating it like it does is where this technique gets misread.