Tools

Why visible results change what a reader can check

A tool that lets you see other submissions' breakdowns lets you check its axes, its spread and its consistency; a tool that hides all results asks for faith.

By Updated 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

Visible results let a reader check a rating tool's spread, whether its axes move independently, and how consistent it is, instead of taking one private number on faith. Public entries turn every other user's submission into a data point you can look at too.

What becomes checkable

Spread. With one result you know nothing about the shape of the scale. With a page of them you can see the range, where the middle sits, and whether the top of the scale is compressed the way most scales are - all of it visible before you submit anything yourself.

Axis independence. If a tool reports a breakdown, public entries let you check whether the axes actually move independently of each other, or whether they rise and fall together as one judgement wearing several labels. That check needs more than one result to run at all - your own submission alone cannot show it.

Consistency signals. Similar-looking submissions that land at wildly different points, or a cluster of entries that all sit suspiciously close together, are both visible in a public feed and invisible in a private one. Neither proves anything on its own, but both are the kind of thing worth noticing before you decide how much weight to put on your own number.

Presentation habits at scale. How rounding, colour thresholds and word labels behave across many results tells you more about the display logic than any single result can, since one result gives you one point on a curve you cannot see the rest of.

What public entries do not settle

Visibility is not calibration. A tool can show every entry and still not disclose how its scale was set, and a page of public results does not tell you what population they were drawn from or over what time period. Machine-learning documentation work asks for more than a feed: the model cards proposed by Mitchell et al. (2019) report "benchmarked evaluation in a variety of conditions" and the context a model is intended for, which no page of entries discloses. It also does not settle accuracy - there is no ground truth on the page to check the scores against, only other scores from the same tool. Measurement has an answer to this that a rating tool cannot borrow: a length taken with a tape and a documented method is checkable against a physical fact, which is exactly the kind of ground truth a rating tool's public feed does not have.

An instance of the property

Rate Cock exposes the breakdown on its public entries, which is the specific property this piece is describing rather than a general endorsement - the axes, not just the total, are visible on entries you did not submit yourself. That is a checkable claim about one tool, and the same check applies to any tool that makes the same claim: look at whether the breakdown is actually there, not whether the marketing says it is.

A quick check you can run in a few minutes

Open a tool's public feed, if it has one, and scroll through twenty or thirty entries. Note whether the totals cluster tightly near the top of the scale or actually spread out - a feed where almost everything sits between 7 and 9 is telling you the scale is compressed at the top before you have submitted anything yourself. If a breakdown is shown, pick two entries with the same total and compare their axes: identical breakdowns behind the same total suggest the axes are not independent, while different breakdowns behind the same total suggest the total really is combining separate judgements. Neither result is a verdict on its own, but both are the kind of thing a private feed simply does not let you check.

What a private tool asks of you instead

A tool with no public feed is not necessarily hiding anything - some have real privacy reasons for keeping every result behind an account - but it is asking you to trust the shape of its scale rather than see it. That trade is closer to what you make with a human reviewer, where a single verdict from one person is the whole of what you get, and there is no larger sample to check it against either. The underlying model behind a private tool is the same kind of system as any other - visibility is a product decision layered on top, not a property of the model itself. A public feed is also only the outside half of an audit: Raji et al. (2020) describe internal audits that produce documentation at each development stage, which a reader of entries never sees. If you want to run this check yourself against two tools rather than take either one's word for it, comparing the two directly is the more direct route, and public entries are simply one more thing you can compare between them.

Read next

Full archive