Tools
Where the rating tool category came from
The ten-point scale, the share card and the crowded middle all predate the current tools; the category inherited its shape before it inherited its models.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Today's AI raters inherited their shape - the ten-point score, the share card, the crowded middle - from two decades of vote and quiz sites that ran no vision model at all. The category took its format first and its computation later.
Stage one: public rating, no model at all
The earliest version of this category was two photos and a vote, aggregated across strangers. No rubric, no axes, no computed anything - just a running average of human clicks presented as a score out of ten. What survived from that era is the presentation, not the mechanism: a single number out of ten, displayed with confidence, next to a photo, built for sharing. The scale itself came from nowhere more rigorous than "ten is a familiar number", and that inheritance is still visible in every tool using it today. What the face-rating branch of this category worked out first - the scale, the share card, the crowded middle - is the direct ancestor of the conventions every later branch, this one included, kept using.
Stage two: the quiz-site interregnum
Before tools could look at an image and produce a judgement, they scored questionnaires instead. A user answered a series of questions about themselves, and the number came from the answers, not from anything visual. This stage did not last as the primary mechanism, but it fixed a pattern that outlived it: a number out of ten, arriving with a paragraph of generated prose explaining it, delivered on a results page built to be captured and shared. The prose-plus-number format did not originate with image models. It was already the house style before there was an image to look at.
Stage three: app-era single-number raters
Smartphone cameras made photo submission trivial, and a wave of apps applied increasingly capable computer vision to the same format the quiz sites had established. One photo in, one number out, still framed by the same share-card conventions. What changed was the input; what stayed constant was almost everything about the output. This is the stage where a single total, with no breakdown, became the default expectation for what a rating tool gives you.
Stage four: breakdowns and multi-axis tools
Only later did tools start reporting components rather than one blended figure. That shift happened for product reasons as much as rigour ones, and it is the first point in the lineage where a tool's output became something you could audit rather than only feel. The current generation of hosted vision models made per-axis scoring cheap enough to offer by default, but the decision to expose it, rather than fold it back into one number, was still a choice each tool made separately.
The crowded middle, inherited too
One more piece of furniture survived all four stages: most results land somewhere in a narrow band above the midpoint, rarely near either end. That pattern started in stage one, where a running average of votes naturally clusters unless the population being voted on is unusually polarised, and clicking voters were not. Stage two's questionnaires inherited it because the scoring rules were written to avoid handing out the lowest values, for the same reason a tool avoids them today - a harsh result loses a user, and a generous one keeps them. By the time stage three's apps and stage four's model-backed tools arrived, a crowded middle was simply what a rating tool's output was expected to look like, and nothing about switching from click-votes to vision models gave anyone a reason to widen it. It is worth noticing that this is a presentation habit carried across four different underlying mechanisms, not a fact about how these subjects are actually distributed.
Why the order of the stages matters
Knowing which piece came from which stage changes what you should expect a "modern" feature to fix. A tool built on the newest model still inherits a scale, a format and a middle designed for older mechanisms, so switching models does not automatically widen the scale, change the format, or thin out the middle - none of those three things are downstream of which model a tool runs. Improving any of them requires a tool to deliberately break with convention, not just upgrade its computation.
What the lineage explains about today's tools
None of the four stages fully replaced the one before it. The ten-point scale from stage one, the number-plus-prose format from stage two, and the single-total default from stage three are all still live conventions inside tools built on stage-four technology. What a current model is actually computing is a genuinely new layer; the scale it reports through, the paragraph wrapped around it and the crowded middle it produces are inherited furniture, not decisions this generation of tools made fresh. Even the model layer tends to learn from stage-one material, human ratings: the research benchmark SCUT-FBP5500 (Liang and colleagues, 2018) labels 5,500 frontal faces with beauty scores from 1 to 5, so that a model can learn an assessment "consistent to human perception."
The lineage also explains what the category never absorbed from its neighbours. It never became a measurement discipline - that stayed with tools built around a tape and a protocol, a different lineage entirely. And it never replaced a human reviewer, who was never part of this chain in the first place and still answers a different question than any of these four stages ever tried to. Rate Cock sits at the current end of this lineage, reporting six axes on a scale whose ten points and whose habit of a running total both trace back further than the model computing them.