Tools
How breakdowns became a feature
Breakdowns arrived when single numbers stopped being enough to keep people coming back; the reason was product, and the effect was auditability.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Rating tools went multi-axis once the single total had run out of novelty, and mainly for engagement rather than accuracy. A breakdown gives a user something to investigate and a reason to return; one number gives them nothing to do after the first look. Transparency was a real side effect, not the motive.
The problem a total could not solve
A tool that returns one number gives a user exactly one thing to do with it: look at it once, maybe share it, and stop. There is nothing to investigate, nothing that changes between one submission and the next except the top-line figure itself. The single total inherited from earlier in the category's history was efficient but had a ceiling - once a user had their number, the tool had nothing left to offer them.
Breakdowns solved that by giving a total something underneath it. A user with a 6.8 broken into three or four components has three or four things to think about, three or four things that might move on a retake, and a reason to come back and check whether any of them changed. That is a product answer to a retention problem, arrived at before it was framed as a transparency improvement.
What breakdowns made possible for readers
The retention motive does not make the effect fake. Once a tool exposes components, a reader can do things a total-only tool never allowed: attribute a change to a specific axis rather than guessing at the whole, notice when every axis moves together and the breakdown was never independent in the first place, or check whether a tool's stated rubric matches what its axes actually do. None of that was available when the only output was one number, regardless of how that number was computed underneath.
The auditability is a genuine gain. It just was not the reason breakdowns got built.
What breakdowns also introduced
Not every axis a tool added was independent of the others. Human raters have the same weakness, and it has a name: "Judgments of character traits tend to be overcorrelated, a bias known as the halo effect," as Westbury and King (2024) put it in a 2024 study testing a new explanation for the bias. A tool can present six labelled scores that all move together, computed from the same underlying signal split into decorative categories after the fact - the version of a breakdown that looks real from the results page and is not. Because the incentive for adding axes was engagement rather than rigour, there was never a guarantee that a given breakdown does what it appears to do; the honest ones and the decorative ones look identical from the results page.
Why the timing lines up with retention pressure, not model capability
It is worth being specific about what actually changed when breakdowns arrived, because it was not primarily a jump in what the underlying model could compute. Simple per-region or per-feature scoring was technically available well before most tools chose to expose it - the computation was cheap even on earlier models, once a tool had already built the pipeline to produce a total. What shifted was a product calculation: at some point, enough of the category had converged on a single-total format that differentiating on "we show you more" became a viable way to stand out, and a wave of tools added breakdowns within a short span of each other for that reason rather than because a new model made it newly possible. That timing is itself evidence for the retention explanation over a purely technical one - if breakdowns had been gated by model capability, the tools would have added them as their models improved, not roughly together once the total-only format had saturated its own novelty.
The failure mode this created
A tool competing on "more axes than the last tool" has an incentive to keep adding categories even after it runs out of things that are genuinely independent to measure. Six labelled axes sound more rigorous than three, whether or not the fourth, fifth and sixth are actually contributing separate information or just restating the first three under new names. Whether a given axis count reflects real independence or diminishing returns dressed up as detail is a question the retention-driven history of this feature makes more relevant, not less - a tool that added axes to compete has no particular reason to have stopped at the number that was actually justified.
Where this shift did not reach
None of this changed how a length gets measured - a tape does not gain axes, because it was never a single blended figure to begin with. It also did not change what a human reviewer gives you, which was always closer to a breakdown by nature - a person naturally comments on more than one thing without needing a product decision to unlock it. The underlying vision models that made per-axis scoring computationally cheap enabled this shift technically, but the decision to expose the breakdown rather than fold it back into one total was a separate, later choice each tool made on its own.
The result today is a category split between tools still reporting one total and tools reporting several, with the split explained more by when each tool was built than by any principled disagreement about what a rating should show. Weighing a newer tool against an older one on this basis alone says more about which side of the breakdown wave each was built on than about which one is actually more accurate. Rate Cock reports six axes and shows the chart on public entries, which puts it on the transparent side of a divide that, on the evidence above, was never really about transparency to begin with.