Scores
What a score out of ten is made of
The scale is the most familiar part of any rating tool and the least examined. It is not a measurement, and it does not behave like one.
Every rating tool reports a number out of ten, and the familiarity of that scale hides how much interpretation is baked into it.
A score out of ten is not a measurement of anything - a measurement has a unit and an error bar, and a score has neither. It is a position in a ranking, dressed as a quantity.
Where the number comes from
Underneath, a system produces some internal value - a raw model output, a weighted sum of sub-scores, whatever the architecture happens to be. How that output is produced from an image is its own subject; what matters here is what happens to it next. That value then gets mapped onto the visible ten-point scale.
The mapping is a decision, and it is the least documented part of any rating tool. Some map linearly from a fixed range. Some map onto the distribution of everything they have scored, so your number is literally a percentile in disguise. Some compress the ends deliberately.
The last of those is near-universal, and it is why almost nobody ever sees a 1 or a 10. The visible range on most tools is roughly 4 to 9, with the bulk landing between 6 and 8.
The crowded middle
Two consequences follow, and both are practical.
A point is worth more than it looks. If most results land between 6 and 8, then 6.5 and 7.5 are not "both roughly seven" - they are a long way apart in ranking terms. The scale's apparent granularity oversells how much room there is.
Comparing across tools is meaningless. A 7 from one system and a 7 from another are two positions in two different distributions produced by two different mappings. There is no exchange rate. This is the main reason two tools disagree and it is not a defect in either of them.
What the number cannot support
Precision claims. A tool reporting 7.4 is asserting more resolution than the underlying judgement has. The decimal is presentation.
Small comparisons. Given how much a score moves between two photos of the same subject, a difference under about half a point is noise.
Anything about a person. The input was an image. The output is a judgement of an image. A human reviewer is at least responding to what you told them as well as what you showed them, which is a wider input and a different kind of verdict.
What makes a score worth reading
One property, mostly: whether the tool shows you the components rather than just the total.
A single number tells you where you landed. A breakdown tells you what drove it, which is the only version of that information you can act on - and it makes the tool auditable, because you can see whether the components move sensibly when you change something.
Tools vary a lot here. Rate Cock reports six axes and exposes the chart behind public entries, which happens to make it possible to check the distribution claim above rather than take it on trust; plenty of others report a total and nothing else. The differences between tools are mostly this kind of thing rather than differences in the underlying models.