Tools
Whether the paid tier is kinder
A paid tier should show more, not score higher; if the same submission scores differently on the two tiers, that is a finding.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Paying should not change the score: a paid tier is meant to buy access - more axes, a saved history, a longer write-up - not a higher number for the same submission. Whether a given tool keeps to that is a separate question, and you can check it yourself by submitting the identical file on both tiers.
Why this is worth suspecting at all
A free tier exists to convert users into paying ones, and a tool has a direct incentive to make the free experience feel like it is missing something - a locked breakdown, a capped result, a number that looks worse than it needs to. The paid tier is where that incentive reverses: a subscriber who feels like they overpaid for the same verdict they already had is a subscriber who cancels.
Neither side of that incentive requires bad faith to produce a real effect. A tool tuned, even unconsciously, to make its paying tier feel more generous is doing exactly what its retention metrics reward it for doing - the same pressure that shapes a cost model's incentives generally applies with extra force at the free-to-paid seam specifically. Outcomes quietly bent by money are a recognised pattern: the FTC's 2022 report Bringing Dark Patterns to Light describes "comparison shopping sites that claim to be neutral but really rank companies based on compensation". None of this is a property of the model doing the scoring - the underlying vision model has no concept of your subscription tier, which means any gap between tiers is coming from a layer the tool built on top, not from anything inherent to the technology.
How to test it
The cleanest version: submit the identical file on the free tier, then submit it again on the paid tier, close enough in time that nothing about the tool itself has changed in between. Compare the headline number, not just the extra detail the paid tier reveals.
Where a tool does not allow both tiers on the same account, a rank-based comparison works almost as well: submit two or three different files on the free tier and note their order, then run the same set through the paid tier and check whether the order holds. Comparing rank rather than raw number sidesteps any legitimate scale difference between tiers and isolates whether the paid tier is systematically kinder.
What a difference would mean
No difference in the number, more detail on the paid tier. This is the tool behaving as advertised - paying buys access to what was already computed, not a friendlier computation.
A small, one-directional difference, paid always a little higher. Worth flagging. A gap in one consistent direction across several tests is not sampling noise; it suggests the paid path is tuned differently, even if only by a fraction of a point.
A large difference, or one that varies unpredictably between tiers. A genuine finding, and one that undermines the free tier's usefulness as a preview - if the free result does not predict the paid one, the free tier is not showing you a smaller version of the truth, it is showing you a different number with a different truth value.
What this test is not checking
It is not testing whether the paid tier is worth its price - what a subscription buys in features, independent of score is usually the real value proposition, and most of the time that value is genuine: more axes, saved history, a fuller write-up. This test isolates one specific and narrower question - does the number itself move - because that is the one difference a tool has the strongest incentive to hide and the reader has the least reason to expect.
It is also not a claim that any particular tool does this. Most probably do not, if only because a documented gap between tiers is the kind of thing that gets noticed and repeated. The value of the test is that it turns a suspicion into a checkable fact rather than leaving it as background distrust that colours every number regardless of whether it is warranted. A tool that shows the same breakdown on both tiers, only gating extra detail rather than the axes themselves - the way Rate Cock exposes its six-axis breakdown on public entries regardless of tier - has made this particular test easy to run and easy to pass.
Running it alongside the rest of the method
This is one of the five checks worth running on any tool before trusting it - the full method covers rubric, calibration and consistency alongside this one, and the paid-tier test fits naturally after the identical-file reproducibility check, since both use the same submit-twice-and-compare structure. Running that reproducibility test first also tells you the tool's own noise floor, which matters here: a small gap between free and paid results might just be the tool's ordinary variance rather than tier-based generosity, and you cannot tell the two apart without knowing the floor first.
None of this maps onto a human reviewer in any useful way - paying more for a person's attention buys a different kind of response, not a different number on a fixed scale, and the free-versus-paid framing this test relies on does not transfer. It also has nothing to do with the physical measurement side of this category, where better gear changes precision and consistency, not a price-gated verdict - two different markets with two different failure modes worth keeping separate.