Scores
Ten is a convention, not a finding
The ten-point scale is inherited from everywhere else, and it imports assumptions the tools never checked.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
Scores are out of ten because ten was already the default for rating anything, not because anyone tested it for this judgement. The category inherited the scale the way a new building inherits the shape of its lot, and ten steps promise more resolution than most tools actually deliver.
Where ten came from
Ten is the decimal habit made visible: it is the count people default to whenever asked for a scale, because it is the number our counting system is built on and the number every other rating context already used. "Rate it out of ten" was a fully formed idiom for judging attractiveness long before an app could do it, going back to informal photo-rating sites that predate any of the current tools. When rating tools needed a scale, ten was not selected for any property specific to the judgement being made - it was the scale already sitting there, familiar, needing no explanation to a new user. Familiarity is not nothing: when Preston and Colman (2000) had 149 respondents rate the same service elements on scales of 2 to 11 categories, the 10-point scale was the one respondents preferred most, although test-retest reliability tended to decrease for scales with more than 10 categories.
What ten assumes
A ten-point scale implies, just by existing, that the space between a 6 and a 7 is the same size as the space between an 8 and a 9. Nothing about how these tools are built guarantees that. The underlying judgement is continuous and the ten steps are a grid laid over it after the fact, and a grid that is evenly spaced on paper is not automatically evenly spaced in what it is measuring.
It also assumes ten meaningfully distinct levels exist to report. In practice, most tools' real output clusters in a narrow band - results bunch between five and eight and rarely touch the outer thirds of the scale - which means the nominal ten steps are doing the descriptive work of three or four. The rest of the scale exists mostly as headroom nobody uses.
What ten costs
The direct cost is a crowded middle: with most results landing in a three or four point range, a scale that looks like it has ten levels of resolution is actually offering something closer to four, and a reader who trusts the nominal precision is trusting resolution the scale is not delivering. The other cost is implied linearity - a reader treats an 8 as meaningfully further from a 6 than a 7 is, because the digits say so, when the underlying calibration may not support that claim at all.
What choosing differently would have changed
A coarser scale would have made this less deceptive by construction, simply by having fewer steps to imply precision with - that comparison is worth its own look rather than folded in here. A scale anchored to something external - a fixed rubric maximum, say, rather than a moving population - would answer a different question than "where do you sit among people who used this tool," and most ten-point scales in this category answer the population question whether or not they say so. None of that makes ten wrong exactly. It makes ten a convention doing a job it was never specifically designed for, on a category that borrowed it because everything else it resembled already had.
The convention shows up differently elsewhere
The same borrowed scale shows up wherever a number is expected, including on tools whose underlying process has nothing to do with the photo-rating sites that made ten the default - what the model underneath is actually estimating is unrelated to why its output gets squeezed onto ten points rather than any other range. Measurement escapes this particular problem because it starts from a unit rather than a convention: a length in centimetres is not competing with any inherited scale, it is just a number with a unit attached. Human review sidesteps it too, in the opposite direction - a person responding in words rather than digits was never forced onto ten points to begin with, and their response is not stretched or compressed by a grid that was chosen for reasons that had nothing to do with them. On a tool like Rate Cock, the total still lands on ten, but the six axes behind it at least let a reader see whether the crowding is uniform across the breakdown or concentrated in one component - which is more than the bare number on its own would show.