Scores

Bell, skew, or pile-up: what the histogram of a tool looks like

Two tools with the same average can distribute their scores very differently, and the shape decides what a given number is worth.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

The shape of a rating tool's score distribution decides how much any single result tells you. Two tools can both centre on 7.2 and still be completely different instruments, because the mostly invisible spread around the average sets what a one-point gap actually means.

Three shapes worth knowing

Symmetric bell. Results spread evenly on both sides of the centre, tapering off toward both ends. This is the shape a reader intuitively assumes, and the one where the arithmetic distance between two scores comes closest to meaning something consistent regardless of where on the scale it sits.

Left-skewed pile against the ceiling. Results bunch up near the top of the scale with a longer, thinner tail stretching down toward the middle and bottom. This is common in tools running generosity bias - pushing the centre up does not just move the average, it also compresses everything above it into a smaller stretch of the scale. Ceiling effects is the direct consequence: less room at the top, more results trying to fit in it.

Bimodal, from a reject threshold. Two separate humps, usually because the tool applies some kind of pass or fail gate before scoring - a low cluster of results that barely cleared the threshold and a higher cluster of results that scored well within it, with a thin gap between them. A tool that refuses or heavily penalises certain submissions before scoring the rest can produce this shape even when the underlying judgement is smooth, because the gate itself carves the distribution in two.

What each shape changes about a mid or high score

In a symmetric bell, a result a full point above the average is a genuinely uncommon outcome, and the gap is informative in a fairly straightforward way. Percentiles make the point concrete: NIST's Engineering Statistics Handbook defines the pth percentile as a value such that "at most (100p)% of the measurements are less than this value," so the same score sits at a different percentile in every differently shaped population.

In a left-skewed pile, the same one-point gap near the top might separate results that are barely distinguishable, because so much of the population is crammed into that same narrow band - while the identical one-point gap lower on the scale, where results are sparser, can represent a much larger real difference. The shape means the scale is not doing equal work at every point along its length, even though it displays as if it were.

In a bimodal shape, a score sitting in the gap between the two humps is unusual by construction, and a score just above the low hump's centre might be closer in practice to the low cluster than the number alone suggests, which makes a small gain across that boundary worth more attention than the raw digit implies.

How to infer shape without the tool telling you

Few tools publish a histogram outright. Some research systems treat the spread as the output itself: Talebi and Milanfar's NIMA model (2017) predicts "the distribution of human opinion scores" for an image rather than only the mean, a reminder that a lone number throws the shape away. A rough sense is still possible from a sample of public results, if the tool shows any - enough submissions to notice where they cluster and whether there is a visible gap anywhere along the range. Rate Cock showing full breakdowns on public entries is the kind of surface that makes this kind of eyeballing possible at all, rather than guessing from a single result.

A tool's own explanation of its scoring, when one exists, sometimes hints at shape indirectly - language about a pass threshold suggests bimodal, language about most users scoring "well" suggests a left-skewed pile, and a description of how the underlying model works occasionally reveals whether a threshold step exists at all.

What this does not cover

This piece is about the general landscape of shapes a distribution can take. The single most common specific outcome inside a left-skewed pile - a huge share of results landing on exactly one whole number - is its own, narrower thing: why so many results are a seven picks that up on its own.

The same shape question applies wherever a bounded scale meets a real population, whatever is generating the number. A human reviewer's scores pile up unevenly for reasons of their own temperament and incentives, and a distribution built from physical measurement data has a shape determined by biology rather than by a mapping choice - a genuinely different kind of distribution, worth knowing apart from the ones this piece describes. Knowing which shape you are looking at is worth more than knowing the average alone ever was.

Read next

Full archive