Scores
Two properties people conflate
A tool that gives the same photo the same score every time is consistent; whether the score is right is a separate question with no clean answer.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
A consistent rating tool returns the same result for the same input; an accurate one returns the correct result. A tool can have either without the other, and conflating the two is the single most common way readers misjudge a score. Only consistency can be checked from outside.
Four quadrants
Consistent and accurate. The tool returns the same number for the same submission, and that number reflects something real about the rubric it claims to apply. This is the quadrant every tool implies it is in and almost none can demonstrate they are in, because accuracy has no independent standard to check against.
Consistent and inaccurate. The tool returns the same number every time, reliably - and that number is simply wrong relative to whatever the rubric claims to measure, perhaps because the scale is miscalibrated or the weighting is broken. This tool is easy to mistake for a good one, because repeatability feels like reliability.
Inconsistent and accurate. The tool's central tendency is right, but any single result wanders enough around it that one draw is not representative. A single submission from this tool is noisy even though the underlying judgement, averaged, is sound.
Inconsistent and inaccurate. The tool wanders and the centre it wanders around is wrong. The worst quadrant, and unfortunately not distinguishable from the second on a single use.
Why consistency is checkable and accuracy mostly is not
Consistency has a direct test: submit the identical file twice and look at the gap. If the gap is small, the tool is consistent on that submission. Run it a few more times, on a few more submissions, and you have a real estimate of the tool's noise floor. This test needs nothing but a second submission and costs nothing but the tool's normal price.
Measurement science draws the same line. Koo and Li (2016), in a guideline on reliability statistics, define reliability as "the extent to which measurements can be replicated", and grade an intraclass correlation below 0.5 as poor and above 0.9 as excellent - a scale for consistency, with nothing in it about being right.
Accuracy has no equivalent test, because it requires a ground truth to compare against, and there is no ground truth for most of what these tools claim to judge. "Objective" is not a property a subjective rubric can have - there is no external fact of the matter that "proportion, symmetry, presentation" is measuring, only a rubric someone wrote down. A tool can be internally consistent about applying that rubric without the rubric itself being correct in any deeper sense, because there is nothing for it to be correct or incorrect against.
This is not the same question as what "calibrated" means for a tool, which is about whether the tool's numbers mean the same thing across submissions and over time. Calibration and consistency overlap but are not identical: a tool can be consistent on a single submission repeated immediately, and still drift in what its scale means six months later after a rubric revision.
What this means for a reader
You can verify consistency yourself, cheaply, today. You cannot verify accuracy at all, and no amount of retesting will get you there, because the thing accuracy would be measured against does not exist independently of the tool's own rubric.
That asymmetry sets a hard ceiling on how much confidence any reader can extract from a rating tool. The honest position is: I can tell you whether this tool is consistent. I cannot tell you whether it is right, and neither can the tool. A tool that at least publishes enough information to make the first check possible - a stated scale, visible spread, an invitation to resubmit - has done what can actually be done. A tool that hides even that is asking for trust on a claim nobody, including its own operators, can verify from outside.
This distinction is specific to automated tools scoring the same rubric twice. A human reviewer is a different kind of variable entirely - asking whether a person is "consistent" means something closer to whether their judgement is stable across submissions, which is its own subject - and a physical measurement sidesteps the whole question, since a tape read correctly is accurate by construction, with only the reading itself to get consistently right. Whatever the underlying model computes, the consistency-versus-accuracy split lives above it, at the scale and rubric layer rather than inside the shared computation most tools run on. Rate Cock publishes six axes per public entry, which at minimum lets a reader test its consistency directly rather than trusting a single number's stability on faith.