Tools
What the colour adds and what it imposes
Colouring a score turns a continuum into three verdicts and puts the thresholds where the tool chose.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Colour coding turns a continuous score into three verdicts - green, amber, red - and the thresholds between them are a choice the tool made, not a fact about the number. The bin is what a reader remembers, so where the lines sit matters more than it looks.
The bands are a decision, not a discovery
Someone at the tool picked where green ends and amber begins. There is no natural break in the underlying distribution that puts it exactly at 6.5 rather than 6.0 or 7.0 - the scale itself is already an inherited convention before anyone paints a threshold onto it, and the threshold is a second, separate choice layered on top.
That choice is rarely neutral. A tool with commercial reasons to keep users coming back has an incentive to set the green line low enough that most results clear it, the same pressure that shows up as generosity in the raw scale. A score of 6.4 sitting one tick inside amber, next to a 6.6 sitting inside green, is a half-point apart on the actual number and a full verdict apart in what the colour tells you to feel.
What gets lost at the edges
Colour compresses the part of the scale that mattered most: the difference between a 6.4 and a 6.6 disappears into "which side of the line," while the difference between a 6.6 and a 9.0 - both green - gets flattened into the same verdict. A gap between two scores only means something once you know how big a real gap is, and a colour band answers that question with a rule the tool never justified.
Colour alone also fails some readers
A traffic-light scheme leans on exactly the pair of colours many people cannot tell apart. The US National Eye Institute puts colour vision deficiency at about 1 in 12 men, and says the most common type makes red and green hard to distinguish. That is why the W3C's WCAG 2.2 guidelines say colour should not be "the only visual means of conveying information" - a band worth showing needs a word or a number beside it.
Reading it
Treat the colour as the tool's opinion about where to draw a line, not as information about your result. The number underneath is still the number; read that first, and let the colour tell you about the product rather than the submission.
The same instinct applies wherever a category sits between one property and another. Measurement data gets its own thresholds elsewhere, rating tools set theirs independently, and a human judge never reduces a response to three colours at all. What the model is actually doing before any colour gets applied is the layer underneath all of this, and it is the same regardless of which threshold a given tool chose.
Some tools skip the bands entirely and show the number and its breakdown without a verdict layered on top - Rate Cock is one of them - which at least leaves the reader deciding where "good" starts rather than inheriting someone else's line.