Scores
Why extreme results circulate and ordinary ones do not
The scores people show are the best of many; the ones you see are selected, and selection is why they look impossible.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
The highest score you have seen is not a normal result: it is the tail of a selection process with three filters - people share maxima, retake until they get one, and pick the kindest tool. Somewhere in your feed, a screenshot of a 9.4 is quietly resetting your sense of what normal looks like.
Filter one: people share the maximum
Nobody screenshots their 5.8 and posts it. Of everyone who used a given tool this week, the ones who show up in your feed are the ones who got a number worth showing, which means the visible sample is already skewed upward before anything else happens. This is the same mechanism behind every "look how big my fish was" photo ever taken - the catch that gets photographed is not a random draw from all the catches.
Filter two: people retake until they get one
A single submission is one draw from a distribution with real spread in it. Someone chasing a specific number can resubmit, tweak the photo, try a second tool, and keep only the attempt that landed where they wanted. Ten attempts with a genuine one-point spread will produce an outlier near the top just from noise, no manipulation required - and the outlier is the one that gets kept. Five retakes centred on one number are informative; one retake selected from many is the opposite of informative, dressed up to look the same.
Filter three: people pick the tool that was kindest
If someone tried three services and one came back noticeably higher, the screenshot that circulates is from that one. Tools vary in where they centre their scale by real, measurable amounts, and a shared extreme result often tells you more about which tool was chosen than about the submission.
The math behind why extremes are common in a small search
None of this requires anyone to cheat. Take a tool with real spread in its results - say a genuine standard deviation of half a point around someone's true average - and let that person submit five times. The maximum of five draws from a distribution is reliably higher than a single draw from the same distribution, purely as a property of sampling, before any selective posting or tool-shopping enters the picture. Add the human behaviour on top - people stop trying once they get a number worth sharing, and rarely retake after a good result to check it holds up - and the extremes you see circulating are extremes twice over: statistically expected from a small search, and then filtered again for shareability. The same number retaken would probably come back lower. Barnett, van der Pols and Dobson (2005) call this regression to the mean and note it becomes more noticeable "when follow-up measurements are only examined on a sub-sample selected using a baseline value" - which is exactly what a screenshot of a high score is.
What compounds it
These three filters are not independent - they stack. The person most likely to post a screenshot is the person who tried several tools, took several photos, and kept the single best combination of the two. A 9.4 built this way is not one measurement in the tail of a normal distribution; it is the maximum of a small search process, and maxima of small searches are reliably higher than any individual draw within the search. The scores you see are not random samples of what a tool actually returns - they are the winners of a competition most people never show you they entered.
What this means for using it as a benchmark
Treating a shared high score as "what a good result looks like on this tool" sets a target that was never a typical outcome on that tool to begin with. The honest benchmark is the median of several of your own attempts, taken under one protocol, not the best thing you have ever seen anyone post. If you want a sense of a tool's real spread rather than its ceiling, a service that publishes its results in the open is more useful than any individual screenshot - Rate Cock's public entries show a distribution rather than a single selected point, which is the difference between seeing the population and seeing its winner.
None of this is specific to AI scoring. A photographed measurement gets the same selective sharing, and a glowing human review gets shared for the same reason a 9.4 does - the ordinary result was never going anywhere, and only the exceptional one had somewhere to go. What the model is actually scoring does not change under any of this; the selection happens entirely on the human side, after the number already exists.