Tools
A survey of uncertainty presentation across the category
Some tools show a confidence value, some a range, most nothing; the choice tells you how the tool wants its number read.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Most rating tools hide a score's uncertainty entirely and print a single figure; a minority attach a confidence value, print a range, or flag a retake. Every score carries real uncertainty whether or not the display admits it, and the format a tool picks says as much about the product as about the model.
Nothing at all
Most result pages show a single figure and stop. No range, no confidence value, no note about image quality - a 7.4 that presents exactly as confidently as a hand-measured constant, when in fact it is one draw from a distribution the tool never shows you. This is the default across the category because it is the cleanest UI, and a clean UI is what most tools optimise for over a technically complete one.
A confidence figure
Some tools attach a separate percentage or star rating next to the score, usually describing something about the input - image clarity, angle, resolution - rather than the model's certainty about the judgement itself. What a confidence figure is actually measuring, decoded is worth reading closely if you have one in front of you, because the two are easy to conflate and the box rarely says which one it is. Even a figure that does describe the model's certainty can overstate it: Guo, Pleiss, Sun and Weinberger (2017) found that modern neural networks, "unlike those from a decade ago, are poorly calibrated", so a stated confidence need not match how often the model is right.
A range instead of a point
A minority of tools print something closer to "7.1-7.7" rather than a single decimal. This is the most honest of the common formats, because it states directly that the number is an estimate rather than a fact, and it puts a rough width on that estimate without requiring the reader to go and compute one. Risk-communication research ranks it the same way: reviewing practice across fields, van der Bles and colleagues (2019, Royal Society Open Science) rank a range above verbal qualifiers on a nine-level scale ordered by precision, with "no mention of uncertainty" second from the bottom, and found some evidence that stating uncertainty "does not necessarily affect audiences negatively". Building your own range from repeated attempts is the manual version of what this format is doing automatically, and the two are worth comparing if a tool's stated range looks either suspiciously tight or suspiciously wide.
A "retake suggested" flag
The least common approach: instead of a number and a range, a flag that says the result may be unreliable and invites another attempt, usually triggered by something detectable in the image itself - poor light, an awkward angle, low resolution. This is closer to an admission of low confidence than a quantified one, and it is the format most likely to come from a tool that has actually built input-quality checks into its pipeline rather than bolting a number onto whatever gets uploaded.
Combinations, and why they are rare
A tool could in principle show all three - a point figure, a range, and an input-quality flag - and almost none do, because each additional element competes with the clean, single-number result for attention on the same screen. Product teams tend to treat uncertainty display as a budget rather than a checklist: pick the one format that signals enough rigour to be credible without cluttering the page the score is meant to headline, and skip the rest.
Why disclosure is uneven across a single product
The same tool sometimes shows uncertainty in one place and hides it in another - a range on the main total but none on the individual axes, say, or a confidence badge on a free tier that quietly disappears once a paid breakdown replaces it with a cleaner-looking chart. This is worth watching for on its own: inconsistent disclosure within one product is a stronger signal about incentives than a tool that simply never shows uncertainty anywhere, because it means the capability exists and was chosen to be shown selectively.
What the choice tells you
A tool that shows nothing is optimising for a clean, confident-looking result, which is a legitimate design choice and also the one least useful to a reader trying to calibrate trust. A tool that shows a range or a flag is accepting a less impressive-looking result in exchange for a more honest one, and that trade is not one every product is willing to make. Rate Cock exposes per-axis breakdowns on public entries rather than a single blended confidence figure, which lets a reader infer spread from the axes themselves even without a stated range.
None of these formats replace checking uncertainty yourself. The manual test - submitting the same file twice and reading the gap - works regardless of what the display shows, and it is the only way to know the tool's actual noise floor rather than trusting whatever figure it chose to print, if any.
The category most comfortable stating uncertainty outright sits outside rating tools entirely. A measurement reported with its method attached treats a stated margin as normal practice rather than a rare disclosure, and a human reviewer can qualify a judgement in words in a way no fixed UI element captures. Why most tools sit on similar underlying models is part of why their uncertainty is similar too, even when the display choices in front of it differ this much.