Scores

Working out a tool's weights from its outputs

If a tool shows a breakdown and a total, a few results are enough to tell whether the total is a plain average or something else.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

Two equal-looking breakdowns total differently when a tool weights its axes unequally, and a handful of its results plus a calculator is enough to detect that. You do not need the source code, because a weighted combination rule leaves a signature in its own output, disclosed or not.

The check itself

Take several results that show both a full breakdown and a total. For each, average the visible axes yourself and compare that plain average to the printed total. If the two match closely across every result, the tool is running something close to an equal-weighted average, and the breakdown is genuinely explaining the total rather than just decorating it.

What a mismatch tells you

A gap that appears once could be rounding. A gap that appears in the same direction across five or six results is a weighting, not noise - some axis is pulling harder on the total than an equal share would predict. Composite-score methodology is candid about why that matters: the OECD and European Commission Joint Research Centre handbook on composite indicators (2008) warns that weights "can have a significant effect on the overall composite indicator," and that whatever the method, "weights are essentially value judgements."

The pattern of the gap tells you more than its size. If the total consistently sits closer to whichever axis was lowest, the combination is minimum-dominated rather than averaged, and the tool is effectively saying that unevenness costs more than a low mean does. If the total tracks one particular axis - say, proportion - more tightly than the others across every result you check, that axis is probably double-weighted or otherwise favoured, even if the tool has never named it as more important. Either way, you have recovered a piece of the rule without ever seeing the code that implements it.

When the total is not derived from the axes at all

There is a case the mismatch check alone will not catch cleanly: a total computed from something other than the shown axes. Some tools generate a headline number from the full internal model output and generate the breakdown separately, as an approximation meant to look explanatory rather than to actually reconstruct the total. In that case the gap between your plain average and the printed total will not be a small, consistent offset - it will vary result to result with no stable pattern, because the two numbers were never mathematically linked in the first place. A wandering gap, rather than a steady one, is the signal that the breakdown is decoration rather than derivation - a rubric with no arithmetic connection to its own total is worth naming for what it is, since it is offering the appearance of an audit trail without the substance of one.

Doing this at any scale that matters

Three or four results are usually enough to see whether a gap is stable. More helps, but the return diminishes fast, because you are not trying to recover exact weights - you are trying to tell "equal average" from "something else" and, if it is something else, roughly which axis it favours. Write the axes and the total down for each result as you go, since the pattern is only visible once you have several rows to compare, and a single result never tells you anything about consistency by itself, no matter how odd that one result looks on its own. This is the same instinct that applies to any calibration a tool declines to state outright: you cannot see the mapping directly, but you can see enough of its output to infer the shape of it.

Where to look for this

Rate Cock publishes both the breakdown and the total on public entries, which is what makes this check possible at all - a tool with no visible axes gives you nothing to compare against, and the weighting question simply cannot be asked from outside. The same transparency gap shows up on the measurement side of the category: a stated method turns a raw figure into something checkable, the same way a visible breakdown turns a total into something checkable, and a number with neither is asking to be trusted rather than verified. A human reviewer has no equivalent weighting to hide in the first place - a person forming an impression is not running a formula over six numbers, so there is nothing here for this particular check to recover, only a different kind of judgement to read on its own terms. And none of this check reaches back into the model producing the axes in the first place - what that model is actually doing to a submitted image is a separate layer, upstream of whatever combination rule the tool applies to its output afterward.

Read next

Full archive