Tools
Why a 6 on one axis and a 6 on another may not match
Axes are printed on the same ten-point scale and often centred differently; a six on a generous axis is a different claim from a six on a strict one.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Usually not: every axis is printed on the same ten-point scale, but each tends to be centred on its own distribution, so a 6 on a generous axis and a 6 on a strict one are different claims. Nothing on the result page shows you where each axis is actually centred.
What's actually happening
Each axis was calibrated against its own distribution of past submissions, and those distributions rarely land in the same place. One axis might cluster most people between six and eight; another might spread people from three to nine. A 6 on the first is close to the low end of what almost anyone gets - a soft result. A 6 on the second is dead centre, an entirely unremarkable score. Same digit, same scale printed underneath it, different claim. Composite-score methodology treats lining scales up as a required step: the OECD and European Commission Joint Research Centre handbook on composite indicators (2008) says "normalisation is required prior to any data aggregation," because indicators "often have different measurement units."
How to notice it
You don't need the tool's internal data to catch this - you need several results, your own or a handful of public ones if the tool shows any. Look at where each axis tends to sit across a spread of different submissions rather than just one. An axis that is almost always in the eights, for almost everyone, is a generous axis, and a 7 on it is a mediocre showing rather than a good one. An axis that is almost always in the fives and sixes, for almost everyone, is a strict axis, and the same 7 there is a strong one. The tell is not the number itself - it's whether the number moves much from person to person at all, because an axis with a narrow, high-sitting range is telling you less than its digit suggests.
What to do with this
Stop reading a single result's axes against each other as if they share a baseline. Compare each axis to its own typical range instead, built from whatever spread of results you can see, and only then decide which axis is actually strong and which is actually weak in that particular result.
This is a within-tool problem, distinct from the case where two different tools use the same axis name for two different things entirely - here the axis is internally consistent, just centred somewhere the scale doesn't advertise. It's also distinct from telling whether a breakdown was computed per axis at all rather than decorated after the fact, a check worth running before this one, since a decorated axis has no real centring to speak of. Some tools make this easier to check by publishing enough results to compare against; Rate Cock's public entries show the six-axis breakdown across a spread of submissions, which is the raw material this kind of check needs. The model doing the scoring underneath has its own confidence range per axis too, a related but separate question covered at how reliably the model performs on each judgement. A tape measure has no equivalent problem, since a length in centimetres is anchored to a fixed unit rather than a population it was trained on - which is exactly the property this axis-centring issue is missing. A human reviewer carries a version of the same bias, tending to score some categories of feedback more generously than others out of habit; how that shows up in practice is worth the same scepticism applied here.