Tools

What a tool that wants to be trusted makes visible

A described scale, a named rubric, a change log, an owner: the things a tool shows when it has nothing to hide.

By Updated 5 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

A rating tool that wants to be trusted shows five things: a described scale, a named rubric, a dated version history, a named owner, and public examples. Each is checkable in under a minute, which makes transparency a short list rather than the general feeling of trustworthiness a well-designed page can fake.

A described scale

Ten points, five stars, a hundred - whichever scale a tool uses, the question is whether it says anything about what the numbers mean beyond their position. Is a 5 the middle of the range, or the middle of the population the tool has scored? Does the scale have a stated distribution, or does the tool simply print a number and let the reader assume it behaves like a percentage? This is a different question from what a measured value like a length would tell you - a recorded measurement comes with its own method and its own disclosure standard, and a rating scale borrows none of that rigour just by printing a decimal.

A tool that states where its average result actually sits is doing more than most: most tools centre somewhere around six or seven rather than five, and a tool that says so upfront is saving the reader from the single most common misreading of any score.

A named rubric

The components a tool claims to judge, listed somewhere with a description of each, not just visible as unlabelled bars on a result page. A named rubric is the difference between "we score six things" and "we score proportion, symmetry, presentation, tone, and two more we have not told you about."

Machine learning already has a template for this kind of disclosure: Mitchell and colleagues (2019) proposed "model cards", short documents that state a model's intended use, "details of the performance evaluation procedures", and how it performs across conditions and groups. The rubric does not need to be exhaustive to count as a signal - it needs to exist as text you can read independent of any one result, so that a surprising score can be checked against a stated claim rather than against a guess at what the tool was even looking at. A rubric also draws the line the tool is claiming to score inside: how the underlying model actually works is a separate, deeper layer most rubrics gesture at without explaining, and a tool honest about that boundary is more useful than one that implies its rubric is the whole story.

A version history for the rubric or the scale

Rubrics get edited and mappings get retuned, and most tools do this silently. A dated changelog - even a short one, even just "adjusted symmetry weighting, March 2026" - is what lets a reader tell whether a shift in their own score reflects a change in them or a change in the tool. Silent change is measurable even in the largest hosted models: Chen, Zaharia and Zou (2023) found GPT-4 identified prime numbers with 84% accuracy in its March 2023 version and 51% in its June version. Without that history, a score from a year ago and a score from today are quietly on two different scales, printed in the same font as if nothing moved.

Ownership

A named operator, a stated company, a page that says who built the tool and what else they run. Ownership disclosure is its own signal worth a full account, and it belongs on this list because a tool that will not say who it is has removed the reader's ability to weigh its interest against its number - every tool has an interest, and the ones that hide the owner are the ones asking to be trusted on faith rather than on the merits.

Public examples

A public board, a gallery of sample results, or - stronger - other users' entries with their breakdowns visible. Public entries are the one feature that actually lets an outsider audit a tool: without them, every claim about the rubric, the calibration or the spread is unverifiable, because the only result you can see is your own.

Rate Cock shows the breakdown on public entries, which means a reader can check whether its six axes actually move independently of each other rather than trusting the claim that they do. That is the kind of property worth looking for - not a rating, a checkable fact about what the page shows.

What the absence of each signal means

None of these five is disqualifying alone. A tool with a described scale but no changelog might be young rather than dishonest. A tool with a rubric but no public examples might be protecting user privacy rather than hiding a thin one.

What is worth noticing is the pattern across all five. A tool missing one signal is probably just under-built; a tool missing all five - no scale description, no rubric, no version history, no named owner, nothing visible beyond your own result - has made every one of its numbers unverifiable at once, and that is a different situation from a tool that simply has not gotten around to publishing a changelog yet.

Where this sits next to the rest of the method

These five signals are one piece of a larger check, not the whole of it - the full ten-minute evaluation also covers consistency and presentation, which transparency alone does not guarantee. A tool can disclose everything on this list and still be inconsistent from one run to the next, the same way a human reviewer discloses who they are without that telling you how they score - disclosure answers who and what, not whether the number itself holds up under repetition.

The five signals here are what a serious tool volunteers before you have to ask. Their absence is not proof of a bad tool, but it is proof the reader has to do more work to trust the one they are looking at, and that work belongs on the tool, not on the user squinting at a results page for clues.

Read next

Full archive