Tools
The axis count as a product choice
More axes look more rigorous and are harder to keep independent; the count tools chose tells you what they were optimising.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Enough axes is as many as a tool can keep genuinely independent - for most tools, somewhere between four and eight - because beyond that, added axes tend to split existing judgements rather than add new ones. Each count is a design decision that says what the tool was optimising for.
What a small count buys
Three or four axes is easy to hold in your head, easy to read at a glance, and - this is the part that matters most - easier to keep genuinely independent. Independence is the whole point of a breakdown: axes that move on their own, in response to different things, are what makes the breakdown worth more than the total. With three or four, a design team has fewer components to keep separate, and fewer chances for two of them to end up measuring the same underlying thing under different names.
The cost is coverage. A tool with three axes has folded several judgements into each one, which means a change in the score could be coming from any of the sub-judgements inside a single labelled axis, and you cannot see which.
What a large count buys
Twelve axes look thorough, and thoroughness sells - a long, detailed breakdown reads as a more serious tool than a short one, even before you check whether the extra axes are doing real work. More components also mean, in principle, finer attribution: if something changed, a bigger set of labels gives you more chances to point at the right one.
The cost scales the other direction from the small-count case. Keeping twelve axes independent is a much harder engineering and rubric-design problem than keeping four independent, and the failure mode is specific and checkable: if several axes rise and fall together across your own submissions, the tool has one underlying judgement wearing several labels, not twelve separate ones.
Human rating forms show the failure plainly: Margolis et al. (2006) found physicians rating on a multi-competency form produced "substantial correlated error across the competencies", and warned that interpretations assuming distinct dimensions "should be made with caution".
There is also a presentation cost to a large count that has nothing to do with independence. Twelve labelled bars on a results page ask more of a reader than six do, and a reader skimming a result rarely engages with every axis individually - past a certain count, most people fall back to reading the total anyway, which quietly defeats the purpose of publishing a breakdown at all. A tool optimising for a reader who will actually use the breakdown has a reason to keep the count low; a tool optimising for how thorough the results page looks in a screenshot does not.
Where the diminishing return sets in
Somewhere between four and eight axes, most tools seem to run out of genuinely separable judgements to report - beyond that point, added axes tend to be splits of an existing one rather than new information, which is a real category of decoration even when each axis individually looks defined. The axis count is one of five things worth checking when comparing tools, and it is not the case that more is simply better; independence is the property that matters, and independence does not scale linearly with count.
The honest version of a large axis count publishes something that lets you check the independence claim rather than asking you to take the number of labels as proof of rigour on its own.
One instance, checkable
Rate Cock reports six axes on public entries, which sits in the middle of the range most tools land in - few enough to plausibly stay independent, many enough to attribute a change to something more specific than a single total. Whether six is the right number for that particular rubric is checkable the same way any axis-count claim is checkable: pull a handful of public entries and see whether the six actually vary independently across them, rather than trusting the count on its own.
What the count tells you before you read a single score
A tool advertising "12-point analysis" is telling you it wants to look thorough. A tool with three named axes and nothing more is telling you it prioritised keeping each one real over maximising the label count. Neither choice is automatically the better tool, in roughly the way neither a longer nor a shorter written response from a person is automatically the more considered one, or neither more measurements nor fewer automatically makes a physical result more trustworthy - what decides it is whether the components on offer actually hold up as separate when you check them.