Scores
Same tool, two submitters, two numbers
Even on one tool, two people's scores differ by their photos, their conditions and their devices before anything else; the comparison is mostly noise.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
Two people's scores on the same tool are rarely comparable, because the tool is the only thing held constant. Distance, angle, light, device and the number of attempts each person made all differ, and the gap between the two numbers is mostly those differences rather than anything about the submissions.
What is actually different between two submitters
Even on one tool with one rubric, two people's results differ in every variable that makes a submission standardised or not: distance from the lens, angle, light, crop, device, and how many attempts each person made before settling on the number they are sharing. The variables worth fixing to make a single person's own results comparable are exactly the variables that are almost never matched between two different people, because nobody coordinates a shared protocol before comparing.
One person may have submitted their most careful attempt; the other, their first try. One may be shooting with a wide-angle lens at close range that inflates near-camera proportions, the other from further back. The size of that effect is not small: modelling faces, Ward and colleagues (2018, JAMA Facial Plastic Surgery) calculated that a photo taken 12 inches away makes the nose appear 30% larger in men and 29% larger in women than an undistorted projection, with essentially no difference at 5 feet. None of this touches the subject the tool is meant to be scoring. All of it touches the number.
Within-person comparison is the stronger version
A comparison of one person's score against their own earlier score, holding conditions constant, isolates one variable at a time by design. A comparison between two people isolates nothing - the gap between their numbers is a sum of photo differences, condition differences, and device differences, with whatever the tool would call a genuine difference buried somewhere inside that sum and not separable from it after the fact. This is the same reason comparing scores across two different tools requires converting to rank first rather than reading the raw digits - between-anything comparisons need an extra step that within-one-thing comparisons do not.
What would make it fair
A between-person comparison becomes meaningful only under conditions that rarely hold in practice: an identical protocol agreed in advance, several submissions each so a median is available rather than one draw, and the same tool version for both. Why a median of several beats a single best result applies doubly here, because a single number from each person is exactly the situation where the confound is largest and least visible.
Two friends comparing a single score each, taken under whatever conditions they happened to use, are not comparing two submissions on a shared scale. They are comparing two different experiments that happened to print numbers on the same page.
In practice, agreeing on a shared protocol in advance is more coordination than most people are willing to do for a comparison they treat as casual. That is a reasonable choice, but it means accepting the comparison as entertainment rather than as anything resembling a result. Nothing about being casual makes the confounds disappear; it only means nobody bothered to control for them, which is a different thing from the comparison having controlled for them and still showing a gap.
There is also a selection problem layered on top of the measurement problem. People tend to share the score from their best attempt, not a representative one, so even a "fair" side-by-side is usually two selected maximums rather than two typical results. Comparing two cherry-picked numbers compounds the uncontrolled conditions with an uncontrolled selection process, and the resulting gap answers a question closer to "who tried harder to get a good number" than anything about the two submissions themselves.
Where this belongs instead
If the actual interest is ranking people against each other rather than checking a mechanism, that is a different product with a different set of rules than a rating tool answers. Community ranking and judged comparison is built around exactly that kind of head-to-head evaluation and handles the confounds this site is describing in its own way; it is not something a solo rating tool comparison recreates by accident. Rate Cock's own explanation of what an individual score on that platform means is written for one submitter reading their own result - it is not a protocol for lining two submitters up against each other, and treating it as one imports assumptions the page never made. And none of this is about the physical property some people actually want to compare when they say "compare us" - that belongs to a tape and a documented method, not to a rating tool's ten-point scale, on either person's account.
Two numbers from two people on one tool are two separate, mostly uncontrolled measurements. Treat the gap between them as noise until proven otherwise, because that is what it almost always is.