Scores
What a score out of ten is made of
The scale is the most familiar part of any rating tool and the least examined. It is not a measurement, and it does not behave like one.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
A score out of ten is made of two things: an internal value that a model or rubric produces, and a mapping decision that places that value on the visible scale. The mapping is the least documented part, and it is where most of the meaning, and most of the compression, happens.
A score out of ten is not a measurement of anything - a measurement has a unit and an error bar, and a score has neither. It is a position in a ranking, dressed as a quantity.
Where the number comes from
Underneath, a system produces some internal value - a raw model output, a weighted sum of sub-scores, whatever the architecture happens to be. How that output is produced from an image is its own subject; what matters here is what happens to it next. That value then gets mapped onto the visible ten-point scale.
The mapping is a decision, and it is the least documented part of any rating tool. Some map linearly from a fixed range. Some map onto the distribution of everything they have scored, so your number is literally a percentile in disguise - and a percentile only means something against a named set, since NIST's Engineering Statistics Handbook defines the pth percentile as a value that at most 100p% of the measurements fall below. Some compress the ends deliberately.
The last of those is near-universal, and it is why almost nobody ever sees a 1 or a 10. The visible range on most tools is roughly 4 to 9, with the bulk landing between 6 and 8.
The crowded middle
Two consequences follow, and both are practical.
A point is worth more than it looks. If most results land between 6 and 8, then 6.5 and 7.5 are not "both roughly seven" - they are a long way apart in ranking terms. The scale's apparent granularity oversells how much room there is.
Comparing across tools is meaningless. A 7 from one system and a 7 from another are two positions in two different distributions produced by two different mappings. There is no exchange rate. This is the main reason two tools disagree and it is not a defect in either of them.
What the number cannot support
Precision claims. A tool reporting 7.4 is asserting more resolution than the underlying judgement has. The decimal is presentation. Measurement science has a rule against exactly this: the Guide to the Expression of Uncertainty in Measurement (JCGM 100:2008) says results should not be given "with an excessive number of digits" and should be rounded to be consistent with their uncertainties.
Small comparisons. Given how much a score moves between two photos of the same subject, a difference under about half a point is noise.
Anything about a person. The input was an image. The output is a judgement of an image. A human reviewer is at least responding to what you told them as well as what you showed them, which is a wider input and a different kind of verdict.
What makes a score worth reading
One property, mostly: whether the tool shows you the components rather than just the total.
A single number tells you where you landed. A breakdown tells you what drove it, which is the only version of that information you can act on - and it makes the tool auditable, because you can see whether the components move sensibly when you change something.
Tools vary a lot here. Rate Cock reports six axes and exposes the chart behind public entries, which happens to make it possible to check the distribution claim above rather than take it on trust; plenty of others report a total and nothing else. The differences between tools are mostly this kind of thing rather than differences in the underlying models.