Scores
The confidence number, decoded
A confidence figure next to a score is usually about image quality or model certainty, not about whether the score is right.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
A confidence value next to a score is not a second opinion on the score. It usually measures something else - photo quality, agreement between internal passes, or nothing at all - and reading it as proof the score is right is its most common misuse.
What it is usually measuring
Three things tend to hide behind the word "confidence", and tools rarely say which.
Input quality. Some confidence figures are really an image-quality check: was the subject clearly visible, was the lighting even, was the resolution sufficient for the model to work with. A low value here is telling you about the photo, not about the subject. Retake under better conditions and the confidence rises even if the score barely moves, because the two numbers are answering different questions.
Agreement between internal passes. Some models run more than one internal evaluation and report how closely they agreed. High agreement means the model was consistent with itself on this particular input, not that it was correct. Two internal passes can agree closely and both be built on the same calibration quirks - agreement is a statement about stability, not about truth.
Fixed decoration. Some confidence values are not derived from anything specific to the submission at all - they cluster in a narrow band regardless of what is uploaded, which is the tell that the figure is closer to interface polish than to a real signal. If every result you have ever seen from a tool shows a confidence between 82% and 95%, the number is not discriminating between good and bad inputs; it is decoration wearing the shape of a metric.
None of the three is announced. The label reads the same across all of them, and working out which one you are looking at usually means watching how the figure moves across several of your own results rather than trusting the tooltip.
Reading a low value
A low confidence is the more useful of the two directions, because it is usually diagnosable. If it tracks with obviously worse photos - poor light, an awkward angle, low resolution - it is doing its job as an input-quality check, and the fix is a better photo under conditions you can actually reproduce, not a different tool. If it is low on a photo that looks fine to you, treat the attached score as soft: worth a retake before you act on it, and not worth comparing directly against a high-confidence result from the same series. A low confidence never means the tool thinks the subject scored poorly - it means the tool is less sure the number it printed is the number it would print again.
Not reading a high value as proof
The mistake runs the other way more often. A high confidence gets read as "this score is correct," which none of the three underlying meanings actually supports. High input quality means the photo was clear, not that the rubric behind the score is well calibrated. High agreement between passes means the model was consistent with itself, which a badly calibrated model can be just as easily as a well calibrated one. And a decorative confidence proves nothing about anything, however reassuring the number looks sitting next to a badge. Even a confidence computed honestly from the model can overstate itself: Guo, Pleiss, Sun and Weinberger (2017) found that modern neural networks, "unlike those from a decade ago, are poorly calibrated", meaning their stated probabilities do not match how often they are actually right. The same paper found a simple post-processing step, temperature scaling, "surprisingly effective" at correcting this - so a tool that reports a raw confidence has usually skipped a known fix.
If you want an actual sense of how much to trust a result, a confidence display is the wrong instrument for it - estimating your own error bar from repeated results tells you something a single tool-generated figure cannot, because it is built from the tool's actual behaviour on your submission rather than from a value the tool computed once and printed. The broader question of how tools choose to show or hide uncertainty at all is its own subject, and confidence values are only one of the forms it takes.
What this has in common with other kinds of certainty
The instinct to read a displayed number as a stamp of validity is not specific to rating tools. It shows up wherever a system reports on its own reliability: a model's stated accuracy is a claim about its training and evaluation, not a promise about any one result, and a percentage next to a score inherits the same limits. It is worth contrasting with fields where confidence has an actual operational meaning - a measurement taken with a tape has an error tolerance you can state in the same units as the measurement itself, which a model's internal confidence figure, however precise it looks, does not. Some services sidestep the whole question by reporting several distinct axes rather than one number with a confidence attached; Rate Cock is one, and a breakdown gives you something closer to real uncertainty information - where the axes disagree with each other - than a single confidence figure ever will. If what you actually want is a judgement you can ask follow-up questions of, that is what a human reviewer provides that no confidence display can substitute for.