Scores
The systematic tilt in rating tools, and where it comes from
A tool that scores low loses users; a tool that scores high keeps them. The centre of the scale moves accordingly.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
Most rating tools score high because a generous number keeps users and a harsh one loses them. Multiplied across a whole user base, that incentive is enough to tilt a tool's scale upward without anyone needing to decide, in so many words, to rig anything.
The incentive
Retention and sharing both reward a generous number. A user who gets a discouraging score is less likely to return, less likely to pay for a repeat, and much less likely to share the result anywhere. A user who gets a flattering one does all three. None of this requires bad faith on a tool's part - it is simply what happens when a product's growth metrics and its scoring output are not walled off from each other, and few products build that wall deliberately.
The pattern is documented well outside this category. Garg and Johari (2018) note that platform ratings "are often highly inflated, and therefore not very informative", and in a randomised trial on an online labour market they found that relabelling the scale with positive-skewed verbal descriptions significantly reduced that inflation.
The mechanism
The generosity does not usually happen at the model's judgement itself - the underlying estimate the model produces is whatever it is. It happens at the mapping step between that internal value and the number that gets printed. A mapping can be shifted upward, or a floor can be applied so results below some threshold rarely surface at all, or the scale can simply be calibrated against a rubric that treats the middle of the underlying distribution as a 7 rather than a 5. Any of these produces the same visible effect: a distribution of results that centres well above the scale's arithmetic midpoint. Where the average actually sits on a tool's scale is the direct consequence of this and worth reading alongside it.
What it does to the reader
The useful property of generosity bias, if it is roughly uniform across the population, is that it inflates every result by something close to the same amount. That means relative reading survives even when absolute reading does not: if your result moved up half a point from your last submission under the same conditions, that movement is still informative, because whatever inflation is baked into the scale was baked in both times. What does not survive is treating the printed number as an absolute claim - an 8 on a generous tool is not the same claim as an 8 on a strict one, and there is no way to tell which kind of tool you are looking at from the number alone.
This is also why a top-end result on a generous tool tends to be compressed - pushing the centre upward means the crowded part of the distribution has to live somewhere, and it is usually the top two points of the scale that end up absorbing it. The most visible symptom of both effects together is a single number that shows up constantly: why so many results land on a seven is the specific case worth reading once the general mechanism here makes sense.
How to tell if a tool is doing this
There is no way to read it off a single result. It takes a sample: a look at a spread of public results if the tool shows any, or a comparison against a tool that is more explicit about its own calibration. A tool that reports per-axis scores rather than one blended total at least lets you see whether the inflation is roughly uniform across axes or concentrated in one - Rate Cock's breakdown-based reporting is a case where the axes are visible enough to check that directly, rather than trusting a single figure.
The same retention pressure exists in adjacent parts of this category that this site does not cover directly: a human reviewer working under a tip or repeat-booking incentive faces a related pull toward kindness, and measurement gear marketed on accuracy has its own commercial reasons to avoid embarrassing a customer. None of that makes any single result untrustworthy on its own. It just means the printed number was never a neutral read of the underlying judgement - it passed through an incentive on the way to the screen, the same way almost every consumer-facing score does once a business depends on the reader coming back for another one, whether that business sells a rating, a review, or a piece of hardware.