Scores
Disagreeing with a rating, productively
Being surprised by a score is not evidence it is wrong; there is a short list of things to rule out before you conclude anything.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
To check whether a rating is wrong, run four checks in order: retake under identical conditions, see which axis carries the surprise, compare your rank rather than the number, and only then try a second tool. Each rules out one specific cause; surprise alone is not evidence of an error.
Retake under the same conditions
The first check is whether the surprise survives a second attempt. Be aware the urge to recheck is lopsided: in Ditto and colleagues' 2003 experiments, people were more likely to spontaneously recheck an unfavourable test result than a favourable one, so apply the same scrutiny to a score that pleased you. Take another photo, matching distance, angle, light and crop as closely as you can, and submit it. If the second number lands close to the first, the result is not a fluke of that single upload - it is what the tool actually thinks, and the disagreement is with the tool's judgement rather than with a glitch. If the second number moves a full point or more with nothing else obviously different, you have found noise rather than an opinion, and the honest next step is a short series, not a verdict on either single result. This only isolates whether the number is stable. It says nothing yet about whether the stable number is a fair one.
Check which axis carries the surprise
If the tool reports a breakdown, look at it before looking at the total again. A total that reads low because one axis is unusually low is a different situation from one where every axis is mildly low - the first points at a specific disagreement you can evaluate on its own terms, the second suggests the whole submission read differently than you expected, which is worth knowing regardless of whether the tool is right. A breakdown exists to be read this way - collapsing it back into "the total is wrong" throws away the one piece of information most likely to explain the surprise. If the tool shows only a total, this step is unavailable to you, which is itself worth noting: a total-only result is harder to argue with because there is nothing to argue with specifically.
Compare rank, not number
If the site shows where your result sits among other submissions, look there next. A number a little lower than expected can still sit in the range you expected, because the scale's centre is not where a school-grade instinct puts it. Rank and score answer different questions, and a surprise that dissolves once you check rank was a surprise about the scale, not about the judgement. A surprise that survives the rank check is a stronger claim, because it means the tool placed you somewhere you did not expect relative to other results, not just printed a number that looked low in isolation.
Try a second tool
Only after the above is a second tool worth reaching for, and only as one more data point rather than a tiebreaker. Two tools rarely share a scale, so the useful comparison is not whether the numbers match but whether the second tool's breakdown, if it has one, agrees on which part of the submission is doing the work. Running that comparison properly takes more care than uploading the same photo twice and eyeballing the results, and it is worth doing right rather than treating one more number as a vote. If what you actually want is a different kind of feedback rather than a second automated opinion, a human judge is answering a different question with a different register, not a more accurate version of the same one.
What none of this establishes
None of these checks can tell you the tool is calibrated well, only whether it is behaving consistently with itself. Consistency and calibration are separate properties: Guo and colleagues (2017) found that modern neural networks tend to be poorly calibrated, producing confidence that does not match how often they are actually right. A rating tool's accuracy claims live with the model that runs it, and no amount of retaking narrows that question - it only tells you whether today's result matches the tool's own pattern. It is also worth separating this from a different kind of surprise entirely: if what surprised you is a length figure, that number never came from a photo in the first place, and the method that actually produces one is a tape and a protocol, not a rating tool at any calibration.
Run the sequence in order and stop as soon as one step explains the surprise. A retake that reproduces the result and a breakdown that shows a specific axis moved is usually enough - you have not proven the tool is right, but you have ruled out the two most common reasons a single result looks wrong. What is left to decide is how much weight to give a rating at all, and that question does not have a single answer that applies to every tool equally.