Photos

A set with different lighting, angles or states in it

A set whose images were taken under different conditions has no single condition; its score describes nothing you can reproduce.

By Updated 2 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

A set is only as good as its worst-held condition, and a mixed-condition set has no condition at all - it has several, unlabelled, folded into one result.

The problem in one line

If one photo in a three-image set was taken level and another was taken from above, or one under window light and another under flash, the set does not represent "the subject under conditions X." It represents a blend of two or three different conditions that the tool cannot separate and the resulting score cannot untangle.

This is a stricter version of the point about multi-photo scoring rules - even a tool that discloses exactly how it combines a set's individual scores cannot rescue a set where the individual scores were never measuring the same thing to begin with. Averaging, best-of, or weighting all operate on the assumption that the inputs are comparable to each other. They are not comparable by default: Hendrycks and Dietterich (2019) built their benchmark of image-model robustness from 15 everyday corruption types, from noise and blur to brightness and contrast, each at five severity levels - the same kind of variation a mixed set carries from one photo to the next. A mixed set breaks that assumption before the combination rule ever runs.

Why you cannot reproduce it later

A comparable set needs defined slots you can refill the same way next time. A mixed set was never built to slots - it was assembled from whichever photos existed, taken under whatever conditions were convenient on whatever days they were taken. There is nothing to write down and repeat, because there was no single protocol behind it in the first place.

The result is a number that cannot be compared to a future set either, because a future set assembled the same loose way will mix a different, unknown combination of conditions. Two uninterpretable results do not become a comparison by being placed next to each other.

What to do instead

Either commit to a single condition per set - the same angle, light and framing for every image in it - or accept that the score is a one-off with no reproducibility claim attached, which is a legitimate thing for a single result to be as long as it is not read as more.

A rating tool that reports per-axis results, the way Rate Cock does on public entries, at least lets you see when a set's axes disagree in ways a consistent condition would not produce - a signal worth reading if a set turns out mixed by accident. The underlying model doing the scoring has no way to flag the mix itself; what it extracts from an image is per-photo, not aware of the rest of the set's history.

A physical measurement sidesteps this entirely, since a tape reading does not depend on lighting or angle at all. A human reviewer, by contrast, will often notice the inconsistency on sight - a judge looking at a mismatched set can say so directly, which is one advantage a person has over a tool that just averages whatever it is given.

Read next

Full archive