Tools
The examples a tool shows you before you submit
A tool's sample results are selected to sell it; they still tell you about the scale, the axes and the tone, if you read them for that and nothing else.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
A rating tool's own demo result reliably shows its scale, axis names, decimal habit, prose tone and presentation, and says nothing about typical scores, accuracy or consistency. The example was chosen, not sampled, so it is worth reading for form rather than outcome.
What a demo result reliably shows
Scale range. Even a hand-picked example has to display a real number on the real scale, so the demo tells you whether the tool runs 1-10, 1-5, or something else, and roughly where the display sits relative to the ends.
Axis names. If the tool reports a breakdown, the demo names the axes it uses. That is useful independent of the number attached to each one - what the axis names actually mean, and whether they overlap with each other, is visible from the labels alone.
Decimal habit. Whether the tool shows a whole number or a decimal, and how many places, is a presentation choice that shows up identically in every result including the sample one. A second digit reads as more precision than it usually carries, and the demo is where you first see whether the tool leans into that or not.
Prose tone. Flattering, neutral or blunt - the demo's write-up is written in the same voice the tool uses for real results, because tone is a product setting applied uniformly, not something dialled up for the sample.
Presentation mechanics. Badge, colour band, animation, share card - all of it appears in a demo exactly as it will in a real result, since building a separate rendering path just for the example would be more work than reusing the real one.
What a demo cannot tell you
Typical scores. A sample result is picked to look good, which means it says nothing about where a typical submission lands on the scale - reading a demo's high number as representative is the most direct way to misjudge what your own result will look like. Advertising law already treats hand-picked outcomes this way: the FTC's Endorsement Guides FAQ says that if an advertiser cannot show a featured result represents what people generally achieve, the ad "must make clear to the audience what the generally expected results" are.
Accuracy. A polished example proves the tool can produce a polished example. It says nothing about whether the underlying judgement is sound, because there is no way to check a demo against anything external - a measurement tool's sample result has the same limit, except the underlying figure could in principle be checked with a tape, which a rating has no equivalent of.
Consistency. One example is one draw. It cannot show you how much the same input would move on a second attempt, which is a property you can only measure with a repeat submission of your own.
What a well-chosen demo tells you by omission
Sometimes the more useful signal is what the demo does not include. A tool that demos a single flattering total without a breakdown is showing you its rubric philosophy as much as its example - a tool built around components tends to lead with them because the breakdown is the selling point, while a total-only tool has nothing structural to show off beyond the number itself. Similarly, a demo that skips the "could not score" or refusal case is not lying, but it is choosing not to show you the tool's edges, and what a tool refuses to score is a real part of its judgement that no polished sample result will ever demonstrate.
Reading the demo without being sold to
Treat the demo the way you would treat a tool's own selected examples in any other category - a curated instance is still real data about form and structure, just not about typical outcome, whether the example is a rating result, a written review, or anything else presented as a showcase. The same caution applies in reverse: if a tool's marketing leans on "AI-powered" language around its demo, that phrase alone describes almost nothing about how the model behind it works, and the demo's actual scale, axes and tone tell you more than the label does.
Rate Cock publishes real public entries alongside its own marketing examples, which is a useful distinction to notice on any tool - a sample chosen by the tool and a public result chosen by nobody are different kinds of evidence, and only the latter lets you see the spread.