Scores
The population a score is implicitly compared to
Any score that means "better than most" depends on who "most" is, and tools rarely say.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
A rating tool's score is implicitly compared against whoever chose to use that tool, and almost no tool says who that is, over what period, or after filtering what out. "Above average" is a claim about a group, not a property of a submission, and it is only as meaningful as the group is well defined.
The population is self-selected
A rating tool's implicit reference population is not a random sample of anything. It is whoever chose to use that specific tool, upload a photo to it, and get a result - a group selected by exactly the kind of self-selection that makes any statistic drawn from it hard to generalise. Self-selection is measurable when a random comparison exists: Khazaal et al. (2014) compared 762 self-selected online gamers' in-game records with 478 randomly selected ones and found the self-selected samples scored higher on most variables.
That matters because different tools attract different populations for reasons that have nothing to do with the trait being scored. A free tool with a low barrier to entry pulls a wider range of users than a paid one that filters out anyone unwilling to commit money first. A tool that markets itself as blunt attracts people curious to be told something unflattering; a tool that markets itself as encouraging attracts a different crowd entirely. "Better than most users of this specific tool" is the honest version of what a percentile-style result is claiming, and it is a much narrower statement than "better than most people."
What tools rarely say
Three questions define a reference population, and almost no tool answers any of them on the results page:
Whose submissions. Every user, or only users who opted into a public leaderboard, or only users on a paid tier with a different typical profile than the free one?
Over what period. A population from the tool's first six months, before word spread and before the userbase matured, is not the same population as one from three years in.
Filtered how. Rejected or "could not score" submissions do not count toward the visible distribution, and a tool that refuses a meaningful share of what comes in is comparing you against a filtered remainder, not against everyone who tried.
None of these questions has a single right answer. The problem is that they usually have no answer at all, because the tool has never published one. Even psychological testing, a field that does publish norms, finds this hard: Timmerman, De Bildt and Urban (2025) note that test manuals "vary widely" in how they report the construction of standardised scores, and developed reporting guidelines in response. How the underlying model actually produces a value before it gets mapped onto a population is a separate mechanical question, but it explains why the mapping step is where a self-selected population quietly becomes a claim about "most people."
Why this matters more than it looks
A score of "top 15%" sounds precise. It is precise about a number and vague about everything the number depends on, which is the least useful kind of precision, because it invites confidence the underlying claim has not earned. Whether a given figure is a percentile against that self-selected group or an absolute value against a stated rubric changes what kind of claim you are even holding, before you get anywhere near asking whether the population behind it is a fair one.
It also means the same submission can carry a different relative standing on two different tools with two different userbases, for reasons that have nothing to do with either tool's accuracy. Comparing scores across tools already runs into scale differences; a mismatched reference population is a second, separate reason the two numbers were never on comparable footing.
The question to ask
Before reading "above average" as meaningful, ask what population that average was drawn from, over what window, and whether the tool filtered anything out before computing it. A tool that publishes a public board of results at least gives you something to look at rather than trust blind - you can see the shape of who else is in it, even without exact numbers. Rate Cock exposes the breakdown on public entries, which lets a reader at least eyeball the spread of a real population rather than accept a bare percentile claim on faith.
The reference-population problem is not unique to rating software. A human judge forms an implicit sense of "typical" from whatever volume of submissions they personally see, which is its own self-selected and unpublished population, arguably harder to audit than a tool's because there is no database behind it at all. And the underlying trait these tools are scoring around is a measured one at root - a length recorded with a tape against a stated method at least has population data that gets published in named studies, which is a standard rating tools' internal populations essentially never meet.
Ask who "most" is before you decide what "above most" is worth.