Photos

The events that retire a reference submission

A new phone, a moved lamp, a tool update, a change of state: any of them ends the series, and the honest move is a new baseline.

By Updated 4 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

A baseline stops being valid when any of four things changes: the device, the environment, the tool, or the subject's state. A baseline is a submission taken under conditions you wrote down, and once those conditions stop holding, every comparison built on it stops meaning anything.

Four events end a baseline, and each one is detectable if you know what to look for.

A device change

A new phone is a new lens, a new processing pipeline and a new colour response, arriving together. Comparability depends on the same device being used every time, not a better one, precisely because a device change moves several variables at once and there is no way to attribute the resulting difference to any single one of them. In metrology terms, a different measuring system moves you out of repeatability and into what the JCGM international vocabulary of metrology calls reproducibility conditions, where the location, operator or measuring system can differ. The tell is simple: if the device in the log does not match the device in your hand, the baseline is already retired, whether or not the numbers have moved yet. A measured length has the same failure mode with the tape rather than the camera - a different instrument, or a different method, produces a different number even when nothing about the subject changed.

An environment change

A lamp gets moved six inches, a room gets repainted, a window gets a new curtain. None of these announce themselves the way a new phone does, which makes them the harder failure to catch. The detection method is the log itself: if the light source you wrote down is no longer the light source in the room, the environment has changed under you. Time of day is a version of the same problem - daylight itself is not a fixed light source, so a series shot at inconsistent hours has an environment change baked into it even when nothing in the room moved.

A tool update

The tool can change without you changing anything. Rubrics get edited, mappings get retuned, and a score from before the change and a score from after it are not on the same scale even though both are labelled out of ten. Hosted models do drift like this: Chen, Zaharia and Zou (2023) found GPT-4's accuracy at identifying prime numbers fell from 84% to 51% between its March and June 2023 versions. This is a fact about the tool's own accuracy claims rather than your protocol, and how a hosted model's read of an image behaves is a subject the tool side owns - what matters here is only that an update on that side quietly invalidates a baseline on yours. Watch for an unexplained jump that affects a fixed control shot as much as a live submission; that pattern points at the tool, not the room.

A change of state

The subject itself can be in a different state between two submissions in ways that have nothing to do with the camera. This is the least mechanical of the four and the easiest to talk yourself out of noticing, because it is also the one you have the least distance on. The honest response is the same as for the other three: if state was part of what the baseline recorded, and state has changed, the baseline no longer describes the current situation. A human judge is not exempt from this either - material sent under one state is not directly comparable to material sent under another, regardless of who or what is doing the looking.

What retiring a baseline actually means

None of these four events mean your history is worthless. It means the series before the event and the series after it are two series, not one, and comparing across the join is the same error as comparing two people's separate scores - technically possible, structurally uninformative. The fix costs almost nothing: log the date of the change, start a fresh baseline under the new conditions, and stop asking the old numbers to explain the new ones. What a stored score actually means already depends on the conditions it was produced under; a retired baseline is just that fact catching up with you a few months later than you expected it to.

This is a narrower question than protocol drift in general. Drift is what happens when standard conditions change gradually and invisibly, session to session; the four events above are discrete and, once you know to look, checkable in a single sitting. Run through them whenever a result surprises you before assuming the tool, or you, did anything wrong.

Read next

Full archive