Tools

The same rating in three wrappers

The same rubric can sit behind a website, an app or a chat bot; the format changes what you see, keep and can compare.

By Updated 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

Web tool, app or chat bot: the format does not change the score, because the same rubric sits underneath all three. What the format decides is what you can see, what gets kept, and whether a later result is comparable to this one.

The web tool

A browser-based tool is usually the fullest presentation: a results page with room for a breakdown, a history if you have an account, and a URL you can return to. That page is also where presentation choices - number, badge, colour, bar, prose - do the most work on how a result gets read, because a web page has the space to use all of them at once. The trade is friction: a browser session, an upload, a wait, none of it especially fast.

The app

A native app usually trims that down in exchange for speed and, often, a saved history by default rather than by request. An app that keeps every result on-device or against an account gives you something a one-off web session frequently does not: a comparable series without you having to build one yourself. That is worth having if you intend to submit more than once and want the tool to support that rather than treat each submission as a standalone event.

What an app usually drops is the breakdown's visual real estate - a phone screen has less room for six axes and a chart than a desktop results page does, so app UIs tend to lead with the total and put the components a tap away, if they show them at all.

The chat bot

A chat interface is the leanest of the three by construction: a photo in, a sentence or two back, sometimes a number embedded in the reply and sometimes not even that. Whatever rubric exists behind it becomes hard to inspect, because a conversational reply does not have a natural slot for six labelled axes - what you get is a summary of a computation you cannot see, formatted as prose rather than as data.

That format also tends to be the weakest on history. A chat log scrolls; it is not built as a comparable record the way an account page with dated entries is, so building a series - the thing that actually makes a repeated score worth anything - is on you to maintain outside the tool rather than something the format hands you.

What does not change

The model layer and the rubric sit underneath all three, and the format is a delivery choice layered on top of them, not a different computation. A tool with a real, named rubric will show it somewhere regardless of format, even if the chat version buries it in a longer reply - Rate Cock surfaces its six axes on the web results page and keeps them on the account history too, which is the format doing the rubric a favour rather than hiding it; a tool with no rubric at all will not gain one by switching wrappers.

That is a contrast worth naming: a tape measurement is method-bound regardless of what records it - the number does not change because you wrote it on paper instead of an app - while a rating tool's output is a function of a model call, and a format that cannot show you the rubric behind that call is asking for more trust than one that can.

What the format does change, reliably, is what you can verify and keep. Metadata handling - what gets retained, what gets stripped - is also format-dependent in ways worth checking separately, since an app with local storage and a chat bot routed through a third-party platform are different privacy shapes even when the underlying model call is identical. Apps sold through Apple's store at least face a written floor: its App Review Guidelines (5.1.1) require every app to link a privacy policy that identifies what data it collects and explains "its data retention/deletion policies", while a web tool or chat bot answers to no equivalent store review.

Choosing on format, not accuracy

None of the three formats makes the underlying score more or less accurate; the model does not know which wrapper called it. What you are actually choosing between is how much of the result you get to see, whether it becomes a record you can compare against later, and how much friction sits between you and a second submission.

If comparability across a series of results matters to you, a format that keeps a dated, structured history is doing real work the others are not - the same discipline that applies to standardising a photo applies to picking a format that will still show you last month's number next to today's, the way a considered human response tends to, one submission read carefully rather than processed in bulk.

Read next

Full archive