Tools
The site's own comparison method, stated
This site puts tools side by side on rubric, calibration, presentation, cost and disclosure; here is what each comparison looks at and why.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
This site compares rating tools on five fixed dimensions - rubric, calibration, presentation, cost model and disclosure - in the same order every time, and never names a winner. Naming the method up front is the only honest way to run a site that compares tools while its publisher also operates one.
The five dimensions
Rubric. Does the tool report components or a single total, and are the components independently checkable. The rubric is usually the single biggest differentiator between two tools, so it comes first.
Calibration. How the internal value maps onto the visible number, and whether that mapping is disclosed anywhere. This is almost never published, which makes it the hardest dimension to compare directly - most of what this site can say about calibration is inferred from behaviour rather than read off a page.
Presentation. Number, badge, colour, prose, animation. Presentation is easy to compare because it is the only dimension every tool shows in full, on every result - it is also the dimension where a tool's own accuracy claims do the least work, since presentation is downstream of the score rather than part of computing it.
Cost model. What a submission costs, what happens on a failed generation, and whether pricing is visible before upload. This is a comparison of policy, not of value for money - two tools can charge the same and handle a refund completely differently.
Disclosure. Whether the tool names its owner, versions its rubric, and shows any distribution information about its own results. A tool that discloses nothing on any of the first four dimensions has, in effect, answered the disclosure question already.
What standardises the comparison
Where a comparison involves submitting the same material to more than one tool, the submission itself is held constant first - same file, same crop, same day. Without that step a difference in tools and a difference in submissions are the same observation, and there is no way to tell them apart afterward.
What "tool" excludes
Two categories sit next to this one and are not part of it. A tape and a method belong to measurement, which produces a length rather than a rating and is compared on a completely different set of dimensions. A person reading a submission and responding belongs to human review, which has no rubric, no calibration and no cost-per-result in the sense this method uses those words - it is compared, if at all, on different terms entirely.
How a single comparison gets built
A comparison piece on this site starts from a specific claim, not a general impression - "this tool shows a breakdown, that one does not" rather than "this tool feels more thorough." Each dimension above is written up only where it is checkable from the outside: a rubric is checkable because you can look at the result page, a cost model is checkable because the pricing is published, a calibration disclosure is checkable because either a distribution exists somewhere or it does not. Where a dimension is not checkable from the outside - the internal weighting behind a total, for instance - the comparison says so rather than guessing, because a guess dressed as a finding is worse than an acknowledged gap.
This also means older comparisons age. A tool's rubric or cost model can change without notice, so a comparison written a year ago is a comparison of that year's version, and the method above is applied fresh each time a tool is revisited rather than carried forward from a prior write-up.
What is not compared, and why
Accuracy. There is no ground truth to measure a rating tool against, so nothing on this site claims one tool is more accurate than another. What can be measured instead is consistency - whether a tool gives the same submission close to the same score twice - and that is what gets reported in its place. That is repeatability in the metrology sense: the International Vocabulary of Metrology (JCGM 200:2012) ties it to the same procedure, the same measuring system and the same conditions, on the same object over a short period of time.
Which tool is "best". A rubric that a reader wants and a rubric that suits someone else are not the same rubric, so "best" collapses into "best for what", and the comparisons here stop at naming the trade-off rather than picking a winner.
The disclosure this method requires
This site's publisher also operates Rate Cock, one of the tools that could appear in any of these comparisons. Where it does appear, it is described the same way as any other instance - a named, checkable property, such as reporting six axes and showing the chart on public entries - never as the winner of a comparison this site would have every reason to rig. That is the entire safeguard: no comparison here declares a winner, so there is nothing for the conflict of interest to bend.
Where to run this yourself
The method above describes how this site compares tools; running the same comparison on two tools of your own choosing follows a shorter, more mechanical version of it - same file, a small set, ranking compared before raw numbers. If what you actually have is one score and a suspicion it is wrong, that is a different question, and checking a single result starts from the score itself rather than from a second tool.