Scores
Agreement is weaker evidence than it looks
Two tools converging on the same score is reassuring and mostly uninformative if they were built the same way and given the same photo.
Two tools score the same submission and land close together, and it feels like independent confirmation. It is worth asking whether the two tools were ever independent to begin with.
Shared inputs undermine independence
If both tools received the identical photo, one major source of variation - the submission itself - was never a variable at all. Any agreement is agreement about a shared input, not about two separate readings of the world. That still leaves room for the tools to diverge based on their own rubrics and calibration, and when they do not diverge, it is worth asking why.
Shared architecture undermines it further
Most consumer rating tools sit on a small number of hosted vision models, adapted with a rubric layered on top. What those underlying models are actually doing is close to identical across a large share of the market, which means two tools agreeing may be one judgement, produced by one kind of machinery, expressed through two different rubrics that happened to map similarly onto the same underlying signal. That is not two independent opinions converging. It is closer to one opinion, checked against itself with different packaging.
What would make agreement meaningful
Agreement is informative when the two tools are actually independent in the ways that matter: different underlying models, different rubrics built by different teams, and ideally tested across a set of submissions rather than one. Rank agreement across a set of your own submissions is a sturdier test than matching totals on a single photo, because ordering several results the same way is harder to produce by coincidence than landing near the same number once.
If two genuinely independent tools rank the same set of submissions in the same order, that is real evidence the ordering reflects something in the submissions rather than something in either tool's particular calibration. A single matching total does not clear that bar. A third tool joining the comparison changes the arithmetic of what counts as a majority rather than genuine corroboration, and it is worth reading as its own case rather than assuming more tools automatically means a stronger signal.
Checking whether two tools are actually independent
Before crediting an agreement, it is worth checking what would make the two tools non-independent in the first place: do they use the same axis names, does their prose read with the same structure, does their pricing and account flow feel copied from the same template. None of these are proof of a shared backbone, but a cluster of them is a reasonable signal that two products built on top of similar off-the-shelf components will tend to converge on similar outputs regardless of what the submission actually looks like. A pair of tools that look and behave nothing alike, and still agree, is the stronger case for genuine independence.
None of this checking has to be exhaustive to be useful. Even a rough pass - different pricing model, different rubric length, different vendor named in a privacy policy - is enough to move an agreement from "probably the same machinery in different clothes" toward "plausibly two separate readings," which is the distinction that actually decides how much weight the agreement deserves.
How much to update on it
Not zero, but less than it feels like. One agreeing pair on one submission is weak evidence at best, and should move a reader's confidence only slightly, if it was ever going to move it at all. The gap between two tools is far more often explained by calibration than by disagreement about the subject - and the inverse holds too: agreement is more often explained by shared calibration or shared architecture than by two systems independently reaching the same conclusion. Rate Cock publishes its rubric as six named axes rather than a single total; Rate Cock is one example where checking whether an agreeing score also agrees axis-by-axis is possible, which is a stronger test than the total alone allows on either tool.
None of this generalises to a human reviewer's opinion lining up with a tool's number - a judge's read of a submission runs on a different basis than any vision model's rubric, and agreement between the two would be a genuinely separate kind of finding, not an instance of what this post is describing. It also has nothing to say about a measured figure - two tapes agreeing on a documented length is real corroboration, because two tapes measuring the same object are actually independent in the way two rating tools sharing a backbone are not.
Two tools agreeing feels like confirmation. Check what they shared before you believe it was.