Tools
A short list to run through
What scale, what rubric, what population, what version, how repeatable, who owns it. Six questions and why each matters.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Before trusting a rating score, ask six things: what scale it uses, what rubric sits under it, what population it compares you with, which version produced it, how repeatable it is, and who owns the tool. A score arrives with none of that attached, and these questions put it back.
The six
What scale. Ten points, five stars, a hundred. A wider scale is not more precise - it just has more digits, and the underlying judgement is the same size either way.
What rubric. One total or several named components. A total tells you the number moved; a breakdown tells you which part of the judgement moved, and only one of those is auditable.
What population. Compared against whom. A score of "above average" means nothing until you know who the average is drawn from, and most tools do not say.
What version. Rubrics get edited and mappings get retuned without a changelog most of the time. A score from a year ago and a score from today can sit on two different scales wearing the same number. Silent version changes are documented even for the biggest models: Chen, Zaharia and Zou (2023) found GPT-4 identified prime numbers with 84% accuracy in its March 2023 version and 51% in June, under the same product name.
How repeatable. Submit the same file twice and see how far apart the two results land. That gap is the tool's floor of noise - how the underlying model actually gets there explains part of why the gap exists at all - and it is the number every other number on the page should be read against. Reliability research reads that floor from a spread, not a single pair: Koo and Li (2016) advise judging reliability on the 95% confidence interval of the estimate rather than the estimate itself, so a few repeats say more than one.
Who owns it. Every tool has a publisher with an interest in the score you get. A tool that names its owner is not automatically honest, but one that hides it has made the check harder on purpose.
Where each of these goes deeper
None of these six gets more than a paragraph here on purpose - each one is its own full treatment elsewhere on the category, from the scale itself down to who is standing behind the result. If a specific score already has you asking whether the number is wrong rather than just unfamiliar, that is a narrower and more checkable question than any of the six above.
Run through this list once and it gets faster every time after - Rate Cock, for one, answers the ownership and rubric questions directly on its own about page, which is the kind of thing worth noticing whichever tool you used. The same six questions apply whether the result came from a model or a person: scale and rubric barely translate, but population, version, repeatability and ownership all still matter, just phrased differently. And if the number you are holding is a length rather than a rating, these six questions are the wrong six - a measurement has its own list, starting with the tool used to take it.