Scores
Calibration, defined for this category
A calibrated tool is one whose numbers mean the same thing across submissions and across time; it says nothing about whether the numbers are right.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
For a rating tool, "calibrated" means its numbers mean the same thing from one submission to the next and from one month to the next. It does not mean accurate. A tool can be consistently wrong and still calibrated, which makes calibration a narrower claim than the marketing use suggests, and a more checkable one.
Consistency of meaning, not accuracy
Picture a tool that is perfectly consistent but wrong in a fixed way: it always reports a value one full point lower than some hypothetical true assessment would. That tool is calibrated - a 7 always means the same thing relative to a 6, and the gap holds steady - and it is also wrong, in the sense that its numbers do not match whatever external standard you might want to check them against. Consistency and accuracy are different claims with no clean way to convert one into the other, and calibration is squarely about the first - how accurate the underlying model actually is is a question calibration alone was never built to answer.
Machine learning uses the word more strictly. Guo et al. (2017) define confidence calibration as predicting probabilities "representative of the true correctness likelihood", and found modern neural networks poorly calibrated out of the box. That stricter sense needs a ground truth to check against, which a subjective rubric does not have, so in this category the usable meaning is the weaker one.
A tool can be calibrated and unhelpful. It can also be inconsistent and occasionally right by accident. Neither pairing is a contradiction, because the two properties are answering different questions, and a reader who conflates them ends up trusting a stable number for the wrong reason or dismissing an unstable one that happened to land close to the truth.
What has to hold fixed
For a tool to stay calibrated, several things need to not move underneath the display without the reader being told.
The rubric - what the tool is actually scoring, and how much each part counts - has to stay the same, or any change has to be dated and disclosed. The mapping from the model's internal judgement to the printed number has to stay the same, or a retuning event has to be visible. If the score is percentile-based rather than absolute, the population it is being compared against has to be reasonably stable, or the tool needs to say when a shift happened. Calibration drift over time is what happens when any of these move without disclosure, and it is the normal state of most tools rather than the exception.
Two checks a reader can actually run
Does the tool describe its scale at all? A tool that states what it measures, how many axes it uses, and roughly what the middle of its range looks like has given you something to hold it to later. A tool that offers only a bare number has given you nothing to check calibration against, which is not the same as the tool being uncalibrated - it just means you cannot tell.
Does it announce changes? A version note, a changelog, any dated marker that the rubric or the mapping moved. This is rare. Its absence does not prove a tool drifts silently, but its presence is a genuine signal that a tool is willing to be checked.
How rating tools differ covers calibration disclosure as one of five things separating tools generally; this piece is about what the word itself is actually claiming, once you strip the marketing use of it away. Rate Cock reporting a named breakdown per axis is the kind of structure that makes a calibration check possible at all - a bare total gives a reader nothing to compare against a later result besides the total itself.
Where the word gets misused elsewhere
"Calibrated" shows up as marketing language across adjacent categories too, and it is worth reading with the same skepticism wherever it appears. A measurement method can be genuinely calibrated against a physical standard in a much stronger sense than any rating tool can claim, because a tape has a fixed unit behind it. A human reviewer calling themselves consistent is making the same narrower claim this piece describes, not a claim about being right - a person can be reliably harsh or reliably generous and still be internally calibrated in the sense that matters here. The word is doing real work in all three cases. It is just doing less work than it sounds like it is doing.