Photos

Matching the protocol to the comparison

A protocol for spotting a two-point difference can be loose; one for spotting a half-point difference has to be strict. Strictness follows the question.

By 4 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

A protocol needs to be as strict as the difference you are trying to see: loosely held conditions can confirm a two-point change, while a half-point change needs every variable fixed. Picking a level of discipline before picking that size gets the effort backwards almost every time.

The question decides the answer

Every uncontrolled variable adds noise, and noise only matters relative to the signal you are hoping to detect. The full reference protocol fixes eight variables at once, and treating all eight as equally mandatory regardless of what you are trying to learn is how people either over-invest in discipline they do not need or under-invest in discipline the comparison actually requires.

A loose case

Suppose the question is whether a submission taken six months apart, after a general change in circumstances, moved by something like two points. A gap that size will usually survive a fair amount of sloppiness in the conditions - a slightly different time of day, a background that is plain but not identical, a device that is the same model but not the same unit. The variables still matter in principle, but a two-point signal is well clear of the amount of noise a loosely held protocol typically introduces, so the marginal value of tightening further is small.

A strict case

Now suppose the question is whether a submission moved by half a point after one specific change - a new light source, say, or a claimed improvement in a single axis. A half-point shift is usually inside the ordinary noise of the tool itself, before any variation in your conditions is added on top. Detecting a gap that small means holding every other variable still enough that its own contribution to the noise floor is smaller than the half-point you are trying to see - same device, same distance to the centimetre you can manage, same hour, same crop, the whole list from the reference protocol treated as non-negotiable rather than best-effort.

Why this is not a call for more precision

Physical measurement uses this exact logic when deciding how much precision a method needs to defend. A documented measuring method is only as strict as the claim it has to support, and over-specifying it wastes effort the same way an under-specified rating protocol wastes trust. The rule here is the same shape: match the discipline to the claim, not to some fixed idea of rigor. Metrology writes the strict end down precisely: the International Vocabulary of Metrology (JCGM 200:2012) defines a repeatability condition as "same measurement procedure, same operators, same measuring system, same operating conditions and same location", with replicate measurements over a short period of time.

Reading the result once the strictness is set

Whether a gap you observe is worth taking seriously is a separate question from how strict the protocol producing it was. How big a gap between two scores has to be before it counts as a real gap, rather than noise repeating, depends partly on the tool and partly on how tightly your own conditions were held - a loosely protocolled two-point gap and a tightly protocolled two-point gap are not equally trustworthy, even though the number on the page is identical.

A middle case

Between those two lives the most common real situation: you noticed a change of maybe a point and want to know whether to believe it. A point is neither trivially large nor vanishingly small relative to ordinary noise, which means it is the range where the strictness of your protocol actually determines the answer you get. Under a loose protocol, a point of movement could plausibly be explained by variation in distance, light or crop alone, and the honest conclusion is that you cannot tell. Under a protocol held to the standard described for the half-point case, the same one-point gap is harder to explain away, and believing it becomes reasonable. This is the practical reason to default toward the stricter end when you are unsure which case you are in: a strict protocol can still detect a large gap, but a loose one cannot rule out a small one.

Applying it

Decide the size of difference you actually care about before you shoot anything. If it is large, a reasonably careful protocol is enough, and chasing perfect consistency past that point is effort spent on precision the comparison does not need. If it is small, treat every line in the protocol as fixed, because a model reading fine differences in an image will register variables you would otherwise dismiss as close enough. A human reviewer applies the same threshold logic from the other direction - a small, deliberate change is worth flagging to them explicitly, because they cannot infer from the image alone how tightly you held everything else. Rate Cock's per-axis breakdown is one of the few places this shows up directly: a half-point move on a single axis is a claim worth checking against a strict protocol before you believe it, and a two-point move rarely needs one.

Read next

Full archive