Photos

Zero variance is also information

If two genuinely different photos return the identical score, the tool may be caching, rounding coarsely or not looking closely; too little noise is a finding too.

By Updated 3 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

Two genuinely different photos returning the exact same score is not proof of a stable tool. It usually means the tool cached a result, rounded coarsely, or is not looking closely enough to tell two inputs apart. The same file resubmitted should match; two separate shots, with normal hand-shake and light flicker between them, should not match exactly.

Why exact agreement is unlikely on its own

Real photographs vary from each other in ways too small to see by eye - a slightly different distance, a slightly different tilt, different noise in the sensor read. A tool that is genuinely evaluating the image should reflect at least some of that variation, even if only in a hundredth of a point. Vision models are known for the opposite problem: Azulay and Weiss (2019) showed that convolutional networks generalise poorly to small shifts and rescalings of an image, and concluded that making them invariant to such changes "remains unsolved". Two different photos landing on the exact same displayed number, repeatedly, is a coincidence you would not expect from a process that is actually looking at pixel content each time.

Three explanations, none of them "the tool is great"

Caching. Some services key a result to something coarser than the full image - a hash of a downsized thumbnail, a rounded set of derived features - and two visually distinct photos can collide on that coarser key. The score you are seeing may not have been recomputed at all the second time.

Coarse rounding. If the tool only displays to a whole point, or rounds an internal value aggressively before showing it, a real but small difference between two photos can vanish in the display layer. The underlying value moved; the number you can see did not.

Not looking closely. A model that is weighting most of its output on a few coarse features - overall brightness, rough proportions, general framing - will be insensitive to the finer differences between two similar-but-not-identical photos, and will return near-identical scores for images that a closer look would separate.

What this changes about how you read a series

A tight spread across a short series is usually the goal - five retakes summarised as a median and a range are meant to converge, and convergence is normally the sign that a series went well. The distinction is whether the spread is tight or zero. Tight and nonzero is what a well-behaved series looks like; exactly zero, across photos that were not literally the same file, is worth a second look rather than a passing grade.

If you want to isolate whether zero variance is coming from the tool's own pipeline or from your submissions being closer to identical than you think, comparing this against the identical-file test separates the two - a tool that also returns identical results on a genuinely repeated file is behaving consistently with itself, while a tool that varies on the identical file but not on two different photos is doing something stranger.

The same suspicion applies outside this site's remit in its own form. A measuring tape either reads a value or it does not, and exact agreement between two readings is the expected outcome, not a red flag - what "the same result twice" is supposed to look like for a physical measurement is a genuinely different situation from a model's output, precisely because a tape is not interpreting anything. A model's insensitivity to small input changes is a property of what the model actually attends to - how much of a hosted vision model's output depends on coarse versus fine image features is the underlying reason a tool can fail to notice two photos are different at all. A human reviewer runs the opposite risk entirely, noticing too much rather than too little - what a person actually responds to across two different submissions rarely lands on the identical read twice, which is its own kind of tell.

Zero variance on non-identical inputs deserves the same scrutiny you would give a suspiciously large jump. On a tool like Rate Cock, where a breakdown is shown per axis rather than only a total, this is easier to check: if every axis, not just the total, matches exactly across two different photos, that is a stronger version of the same finding and worth taking seriously.

Read next

Full archive