Tools

Drawing the boundary of the category

A tool that returns a number is a rater; a tool that returns an edited image or a compliment is something else wearing the name.

By Updated 4 min readTools

Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it

A rating tool is anything that takes a submission and returns a repeatable judgement on a defined scale or rank. Filters, compliment generators and quizzes that reflect your answers back borrow the name without making that claim, and the boundary is worth drawing precisely.

The claim a rating tool has to make

A rating tool takes a submission and returns a repeatable judgement on a defined scale or rank - the same input should produce, allowing for the model's own noise, roughly the same output twice. That is the whole test. Metrology has a precise version of it: the international vocabulary (JCGM VIM, entry 2.20) defines repeatability conditions as the same procedure, operators, measuring system, operating conditions and location, with replicate measurements "over a short period of time". It does not matter whether the output is a decimal, a percentile, or a category, and it does not matter whether the axis is subjective, adjacent to measurement, or something else entirely - what matters is that there is a scale, and the tool is claiming to place the submission on it consistently.

Two consequences follow from that definition. First, the output has to be structured enough to compare against a second attempt - a number, a rank, or a defined category all qualify, a paragraph of free text does not. Second, the tool has to be applying something like a consistent procedure, even an opaque one, rather than generating a fresh, unrelated response each time.

What borrows the name without the claim

Filters. An app that returns an edited image - smoothed, brightened, retouched - is not rating anything. It is producing an image, and the fact that the image looks more flattering than the input tells you about the filter, not about a score the original earned.

Novelty compliment generators. A tool that returns "you're a solid 9, gorgeous!" regardless of what was uploaded is not applying a rubric; it is producing a response designed to be shareable. The way to tell the difference from a genuine rater is the one test that actually distinguishes them: submit two different photos and see whether the output changes, and by how much. A tool where the retake test finds nothing has told you something - either it is not scoring the submission at all, or the scale is so compressed it might as well not be.

Quizzes with no scale. Some apps ask a set of questions and return a result, but the result is a self-report summary rather than a judgement of anything submitted - the input was your answers, not an image, and the output describes what you said rather than assessing it. A genuine quiz-style rater still needs a defined scale behind the questions to count; one that just reflects your answers back in different words does not.

Beauty filters marketed as scores. A near relative of the plain filter: an app that applies a cosmetic edit and then attaches a number to the edited result, as if the number were assessing the original submission rather than its own output. This one is worth calling out separately because it is the most deceptive of the three - it has the surface of a rater, a decimal and everything, while the actual judgement being reported is of an image the tool itself altered first.

Where the boundary gets genuinely fuzzy

Label-only tools sit right on the edge. A tool that returns a category rather than a number - "above average," "elite," a tier name - is still making a repeatable, structured claim, just a coarser one, and it counts by the definition above even without a decimal anywhere in sight.

The harder case is a tool that blends a genuine score with generated flattery in the same output. Rate Cock keeps that boundary clean by reporting its six axes as the actual result and treating the written paragraph as separate, downstream commentary rather than the score itself - a structural choice, not just a style one, and it is what makes the number the thing worth checking rather than the prose around it.

Why the boundary is worth drawing at all

Once you can tell a rater from something wearing its name, the whole rest of the category becomes a question of which shape of rater you are looking at - single-number, multi-axis, pairwise, rank-only, or quiz-style - rather than whether you are looking at one in the first place. That is a different, later question, and it only makes sense to ask once the first one is settled: does this thing claim a repeatable placement on a scale, the way a considered response from a person or a number derived from an actual method both do in their own registers, or is it something that just looks like one on a landing page.

Read next

Full archive