Scores
From 7.2 to "roughly 6.7 to 7.7
A single result is best read as a range; here is one way to build that range from three repeats and a little honesty.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
To turn a score into a range, retake three times under matching conditions, report the middle value, and take half the spread either side of it. A rating tool prints one number, but the honest version of any result is a range, and this arithmetic is simple enough to do in your head.
The three retakes
Say you submit under matching conditions three times and get 7.2, 6.8 and 7.4. Sort them: 6.8, 7.2, 7.4. The middle value, 7.2, is the number worth reporting as your central estimate - the median rather than the highest or an average - and the spread between the lowest and highest, 0.6, is the raw material for a range.
Building the range
Take half the spread either side of the median as a rough working range: 7.2 minus 0.3 and 7.2 plus 0.3 gives roughly 6.9 to 7.5. This is not a formal statistical interval. In the NIST/SEMATECH e-Handbook of Statistical Methods, the interval around a mean scales with the standard deviation divided by the square root of the sample size, with a t-distribution correction for small samples, so a formal interval from three readings would be wide. The half-spread band is a defensible, honest approximation instead: it says "based on what I have actually observed, results land somewhere in this band," which is a truer statement than "my score is 7.2." If the three retakes had been tighter - say 7.1, 7.2, 7.3 - the same method gives roughly 7.05 to 7.35, a narrower band reflecting that the submission is producing more consistent results. The width of the range is information in its own right: a wide range says the result is noisy and any single number from it should be treated lightly, a narrow range says the tool is behaving consistently on this submission, whatever you think of the number itself.
Comparing two ranges
Say a second submission, taken a month later after a change you are curious about, produces retakes of 7.6, 7.9 and 7.5, for a median of 7.6 and a range of roughly 7.4 to 7.8. The two ranges, 6.9-7.5 and 7.4-7.8, overlap only slightly, at the very top of the first and the very bottom of the second. That is the signal worth trusting: a shift from one range to a mostly non-overlapping range is a real difference, in a direction, even without a formal significance test behind it. Compare that with a second submission that produced a median of 7.3 with a range of 7.0 to 7.6 - almost entirely inside the first range - and the honest conclusion is that nothing detectable changed, whatever the two median numbers happen to say on their own. State direction, not false precision: "this looks like an improvement" is a claim the overlapping-ranges method actually supports; "this improved by exactly 0.4" is not, because 0.4 is the gap between two medians drawn from noisy series, not a measured quantity.
What this method does not require
It does not require you to already understand where the underlying uncertainty in a single score comes from - that concept is covered on its own, and this worked example only needs you to accept that repeats vary and to act accordingly. It also does not prescribe how the retakes themselves should be taken; the retake procedure is a separate piece of the method, and this example assumes you already have three comparable numbers in hand.
The same logic elsewhere
Measurement science made this a rule long ago: NIST's guidelines for expressing measurement uncertainty (Taylor and Kuyatt, 1994) set out how a result is reported together with its uncertainty, and advise that "it is preferable to err on the side of providing too much information rather than too little." The instinct to state a range rather than a point estimate is standard practice anywhere a single reading is unreliable on its own, including a measurement taken with a tape, which is reported with a tolerance rather than a bare figure precisely because one reading is not the whole story. A tool that reports several axes rather than one total gives you this for free on each axis separately - Rate Cock shows per-axis figures on public entries, which is a finer-grained version of the same idea, a range per judgement rather than one range for a blended total. None of this substitutes for asking whether the model behind the number is well built in the first place; that question sits with the tool's own accuracy claims, not with how many times you retake it. And if what you actually wanted from the comparison was a specific, worded response rather than two overlapping bands of numbers, a human reviewer gives you that directly, at the cost of not being repeatable the way a rating tool is.