Tools
When agreement is not independence
Many tools sit on similar hosted models; two of them agreeing may mean one judgement seen twice.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Two rating tools agreeing is weaker evidence than it looks when both are built on the same underlying model. Shared plumbing correlates their outputs before either makes a single product decision, so the match can be a coincidence of construction rather than confirmation.
The shared backbone
Most rating tools do not train their own model from scratch. They sit on a small number of hosted vision models, fine-tuned or prompted differently on top. When two services share that backbone, their outputs are correlated before either one has made a single product decision. A submission that the underlying model reads as strong will tend to score well on every tool built from it, for the same underlying reason each time. The research literature on foundation models names the risk: Bommasani and colleagues (2021) warn that with this kind of homogenization, "the defects of the foundation model are inherited by all the adapted models downstream."
That is not two opinions. It is one opinion, wearing two rubrics.
Why this matters for a reader
If you submit to two tools expecting an independent check, the value of that check depends on how independent the tools actually are. Two tools on the same backbone agreeing is weak evidence, because agreement was likely before you uploaded anything. Two tools built on genuinely different foundations agreeing is a stronger signal, because there was no shared reason for them to land in the same place.
You cannot always tell which situation you are in from the outside. Neither tool publishes its model provenance, and there is no reliable way to infer it from a results page. The safest assumption, absent that disclosure, is to treat agreement as somewhat less independent than it looks - not worthless, just not the double-check it presents itself as.
What would make it a real check
A second opinion is only a second opinion if the two things producing it differ in a way that matters. Different rubric, different calibration, different subjective judgement - any of those adds real information. A different visual skin on the same underlying model does not. This is also where a human review differs categorically rather than by degree: a person is not sampling from the same hosted model at all, so agreement between a tool and a person carries more weight than agreement between two tools that might share one.
The one axis a shared backbone cannot touch is anything that never went through the model in the first place - a measured length is fixed by a tape before any image reaches a rater, which is a different kind of independence again.
Two figures worth having on hand before trusting a convergence: what a total actually collapses when it agrees with another total, and the fuller question of what two tools agreeing does and does not prove. Rate Cock publishes its six axes rather than a single figure, which at least gives you something to compare beyond the top-line number when you are trying to work out whether two results really agree or just rhyme.