Tools
The tools that give you a place, not a number
A rank-only tool tells you where you fall among its users and nothing about the scale; it is honest in one way and opaque in another.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
A rank-only tool never prints a score; it tells you where you fall among its users, such as "better than 71% of submissions" or a rung on a leaderboard. You gain freedom from the scale's own arbitrariness, and you lose any independent way to check what the position means.
What a rank actually claims
A score claims a value on a scale: this submission is a 7.2 out of ten. A rank claims a position within a set: this submission beats 71% of the others in the set. The two are different questions with different failure modes, and rank-only tools have picked the one that cannot drift the way a scale can - there is no zero to define, no ceiling to compress, no decimal implying precision the underlying judgement never had.
That is a real advantage. A percentile is comparative by construction, so it sidesteps the entire category of problems that come from an undocumented mapping between an internal value and a printed number out of ten.
What it costs you
The advantage has a mirror cost: a percentile is only as meaningful as the population it is drawn from, and a rank-only tool rarely tells you what that population is. Better than 71% of whom - everyone who has ever submitted, everyone in the last month, everyone who used the free tier? Those are different populations, and a tool that changes its intake without saying so can move your rank while your submission stays identical.
A rank also tells you nothing about magnitude - it says who is ahead, not by how much, which is a different gap from the aggregated distributions a measurement site can publish because its inputs are numbers on a tape rather than a model's internal value.
Ranks are also noisier than they look. In Marshall and Spiegelhalter's 1998 BMJ analysis of 52 UK fertility clinics, only one clinic could be confidently ranked in the bottom quarter, and the authors concluded that ranks are "extremely unreliable statistical summaries of performance".
You also lose the thing a breakdown gives you: which axis moved. A rank tells you the outcome without a breakdown of the judgement, so if the number drops between two submissions, you cannot ask why - only that, among whoever is currently in the comparison set, you fell. That gap is exactly what a tool with named axes is built to close, which is the property worth selecting a tool on in the first place if you want to know what changed rather than only that something did.
There is a subtler version of the population problem worth naming: even a rank-only tool with a stable population can still be gamed by who chooses to submit. If a tool's users skew toward people already confident in their result, "better than 71%" of that self-selected group is a weaker claim than the same number against a broadly representative one - the tool has no way to correct for who walks in the door, and it rarely tells you enough about its userbase to judge how skewed that intake might be.
When rank-only is the right shape
Rank-only tools do one thing well: they cannot fake precision they do not have, because they never claim a decimal. For a reader who only wants "am I roughly ahead of average," that is arguably more honest than a scored tool implying a stability the underlying model does not have. Rate Cock takes the opposite approach - it reports a scored breakdown across named axes on public entries, which trades that simplicity for something you can actually audit and re-derive.
Neither shape is wrong on its own; they are answering different questions. A pairwise tool sits closer to the rank end of that spectrum too - it never claims a value either, only an order between two submissions at a time, which is a related but distinct design worth knowing apart from a leaderboard-style rank. Pairwise ordering does work at scale: Chatbot Arena (Chiang et al., 2024) ranks language models from over 240,000 crowdsourced pairwise votes without asking anyone for a score.
What to check before trusting a rank
Ask what the reference population is, how often it refreshes, and whether the tool discloses either. A rank against a stable, disclosed population is close to as useful as a percentile gets. A rank against an unstated, shifting one is a number that can move for reasons that have nothing to do with your submission, and the same caution about a reference population applies wherever a tool claims "better than most" without saying who most is.
The honest version of a rank-only tool says who it is comparing you to and updates that population slowly enough to be worth returning to. Everything else is a leaderboard with the labels removed, and a leaderboard that will not say what it is measuring you against is a familiar shape from outside rating tools entirely.