Scores

A high total with a low axis, or the reverse

A breakdown that does not match its total is telling you about the weighting; a breakdown that contradicts itself is telling you about the noise.

By 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

When the axes disagree with the total, one of two things is going on. Either the tool weights its axes unequally, which is a fact about its formula, or one axis is noisy, which is a fact about that run. A repeat under the same conditions tells you which, and each deserves a different amount of weight.

Case one: the total does not match the axes

If every axis looks roughly consistent with each other but the total seems too high or too low for what they show, the likely explanation is weighting. A total is rarely a plain average of its axes. Most tools apply some combination of weights, and a weighting scheme can let one strong axis carry the total past what the weaker axes would suggest, or let one weak axis drag it down further than the others deserve. Even a published weighting is not the whole story: Paruolo, Saltelli and Saisana (2011) found, in composite indices including the Human Development Index, that "the declared importance of single indicators and their main effect are very different" in many cases.

This is worth checking rather than assuming. How sub-scores get weighted into one number covers the mechanics - equal weights, fixed but unequal weights, and non-linear combinations all produce a total that reads differently against its own breakdown. A total that looks mismatched against its axes is not necessarily broken; it might just be reporting a weighting scheme you have not seen stated anywhere.

The practical move here is arithmetic, not guessing. If a tool shows both a breakdown and a total across enough results, the pattern of how they relate becomes visible even without a published formula - which is the same procedure that uncovers a hidden weight rather than a plain average sitting behind an apparently simple total.

Case two: the axes contradict each other

The second case is different: not a total that seems off relative to consistent axes, but axes that disagree with one another - one high, one low, with nothing in the submission that obviously explains the split. This is a noise signature, not a weighting signature.

Some axes are inherently more stable than others. An axis tied closely to something structural in the image tends to vary less between repeats than an axis that is closer to a taste judgement, and a genuine contradiction between two axes on the same submission is more likely to reflect one noisy axis than a real, meaningful split in the underlying subject. The error bar a tool never prints applies per axis as much as to the total, and an axis with a wide spread is exactly the kind that will occasionally sit far out of step with its neighbours on a single run.

The way to tell the two cases apart is a repeat. If the same contradiction shows up again on a fresh submission under the same conditions, it is more likely to be a real, structural disagreement between two things the tool measures separately. If it does not repeat, one axis was noisy that time and the disagreement was never meaningful in the first place.

What each case is worth

A weighting mismatch tells you about the tool's formula, which is useful for understanding how to read any of its totals, but not urgent information about the specific submission in front of you. A genuine axis contradiction is more informative: it means the submission is scoring unevenly across dimensions the tool considers separate, which a single total would have quietly blended into one unremarkable-looking number.

A tool that reports six independent axes rather than one blended figure is the only kind where this distinction is even visible to a reader - a total-only tool gives you no breakdown to check for either weighting mismatches or genuine contradictions, and both failure modes just disappear into the single number.

The same two-case logic shows up outside rating tools too. A physical result can look inconsistent the same way - two measurements that do not agree with each other are asking the same weighting-versus-noise question a breakdown is, just with a tape instead of a model. And a written verdict from a human reviewer that seems to praise one quality while marking the overall impression down is worth the same check: is that a stated weighting of what the reviewer cares about, or an inconsistency in how they read the same material twice. Whichever the underlying model is doing when it produces per-axis values in the first place turns out to matter here too, since a model that computes axes independently is more likely to produce genuine contradictions than one deriving them from a single internal judgement split into labels after the fact.

A mismatch is a formula question. A contradiction is a repeatability question. Different fixes, different weight to give the result.

Read next

Full archive