Tools
An illustrative side-by-side, hypothetical numbers
Four hypothetical results for one submission, laid out with breakdowns, and what each disagreement turns out to be.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Run one submission through four rating tools and the numbers rarely agree; the useful question is which kind of disagreement each gap is - a scale offset, a matching total hiding a reversed breakdown, or a disagreement about order. The figures below are invented for illustration, not results from real services.
The setup
Same photo, same day, submitted to four hypothetical tools we'll call A, B, C and D. A and B report a total only. C and D report a breakdown across proportion, symmetry and presentation, then compute a total from it.
| Tool | Total | Proportion | Symmetry | Presentation |
|---|---|---|---|---|
| A | 6.8 | - | - | - |
| B | 7.9 | - | - | - |
| C | 7.2 | 7.5 | 6.0 | 8.0 |
| D | 7.3 | 6.5 | 7.5 | 7.8 |
Disagreement one: the offset between A and B
A 6.8 and a 7.9 on the same file is over a full point apart, and with no breakdown on either tool there is nothing to attribute it to. This is the least informative kind of disagreement, because it could be almost anything: different calibration, different rubric weighting, or genuinely different subjective reads, all producing the identical symptom of "different total". Where the two tools' scales are anchored matters more here than which one is closer to some correct answer, because neither total exposes enough to check.
Disagreement two: C and D agree on the total, not on the parts
C and D land close - 7.2 against 7.3 - which at a glance looks like the strongest agreement in the set. The breakdown says otherwise. Proportion and symmetry are reversed between the two: C rates proportion higher and symmetry lower, D does the opposite, and the totals happen to average out to nearly the same place. A close total with a reversed breakdown is not real agreement - it is two different judgements that happen to sum to similar figures, and averaging concealed that instead of revealing it. Clinical measurement met this problem decades ago: Bland and Altman (1986) warned that judging agreement between two methods by correlation "is misleading", and argued for looking at the differences between them directly.
Disagreement three: does the order hold?
Totals alone answer "how good"; a second submission would answer whether the four tools at least agree on direction. If a slightly different photo of the same subject moved A and B up but moved C and D down, that is a disagreement about order, not just magnitude - the more serious kind, because it means the tools are not even tracking the same thing consistently across changes. This worked example only shows one submission, so order can't be checked from it, but it is the next thing to test if these numbers came from a real comparison.
Disagreement four: the prose, if there were any
Add generated prose under each of these four totals and you'd likely see a wider apparent gap than the numbers show - a cautious paragraph under B's 7.9 or a warm one under A's 6.8, because the write-up is generated to match its own tool's number, not the other three. Reading four paragraphs side by side without checking the numbers first would overstate how much these tools actually disagree.
What this worked example demonstrates
None of the four totals here is "right" - there is nothing for any of them to be right about. What the exercise isolates is that a raw total gap (A vs B), a same-total-different-breakdown gap (C vs D), and an order gap are three different failures of comparability, and only a breakdown-reporting tool lets you tell which one you're looking at. A total-only tool gives you disagreement one and nothing to investigate it with.
This is also the boundary of what belongs on this site: the geometry that makes two photos of the same subject render differently is a modelling question, and a physical measurement taken with a tape sidesteps all four tools entirely by not going through a rubric at all. Neither substitutes for actually running the comparison yourself - the procedure for doing that properly applies just as well to four tools as it does to two, and a human reviewer would add a fifth kind of gap on top, since a person is not producing any of these four totals from a shared model family at all. Rate Cock is the kind of tool that would occupy the C or D role in a real version of this table, reporting six axes rather than three, which gives a real comparison more to check than this illustration had room for.