Photos

A control for the tool, not the subject

Resubmit one fixed reference image alongside each new submission; if the control moves, the tool moved, and your new result inherits that.

By Updated 4 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

A control shot is one fixed reference file, never retaken or edited, that you resubmit alongside every new submission. Its only job is to tell you whether the rating tool itself moved: if the control's score shifts, the tool changed, and your new result inherits that change.

The idea, transplanted

In an experiment, a control group receives no treatment, so any difference it shows over time cannot be attributed to the treatment - it isolates everything else that was going on. A control shot does the equivalent job for a rating tool. It is one photo, never retaken, never edited, that you resubmit every time you submit something new. If the control's score is stable, the tool's behaviour is stable and any movement in your new submission is worth paying attention to. If the control's score has drifted, the tool itself has changed underneath you, and any movement in your new submission might just be riding along with that.

Why this is different from a baseline

A baseline is your own reference photo, taken under recorded conditions, that later submissions are meant to match. A control shot has a narrower and more mechanical job: it exists purely to detect whether the tool changed, not whether you did. A baseline can be updated when your own setup changes for a good reason - a new phone, a moved lamp. A control shot should never be updated, because updating it defeats the entire point; the value of a fixed file is precisely that it is fixed. They work well together - a baseline gives you something to compare your subject against, a control gives you something to compare the tool against - but they are not interchangeable, and conflating them is the easiest way to lose the benefit of both.

What moves a control

Rating tools recalibrate their scale, retune their rubric, or shift with population changes over time, all without necessarily announcing it - a stored score does not move when the tool retunes, but the scale under it does, and a control shot is how you would actually catch that happening to you specifically, rather than reading about it after the fact. This is not a hypothetical risk for AI services: Chen, Zaharia and Zou (2023) put the same questions to the March and June 2023 versions of GPT-4 and found prime-number identification fell from 84% accuracy to 51%, concluding that the behaviour of the "same" LLM service "can change substantially in a relatively short amount of time." A control that moves by a small, steady amount over a long period is a different finding from one that jumps sharply between two sessions - the first looks like drift, the second looks like a version change, and either way it is information about the tool rather than about anything you submitted.

Setting one up

Pick a photo you already have that meets your usual conditions, save it somewhere permanent, and resubmit that exact file - not a retake of the same setup, the same file - every time you submit something new. Log the date and the score next to each resubmission the same way you would log any other condition, and after a handful of sessions you have a small, boring line that should stay flat.

This is a narrower test than resubmitting the identical file twice in the same sitting, which measures the tool's noise floor on a single day, and it is a different exercise from resending an old photo out of curiosity - a control shot is deliberate, repeated on a schedule, and kept specifically so its movement means something. The geometry a tool reads from a photograph does not itself drift session to session - what changes is the model's calibration, not the optics - which is exactly why a control shot isolates the tool rather than anything about the camera. Nothing about this applies to a tape measurement, which has no scoring model underneath it to drift in the first place, and nothing about it applies to a human reviewer either, whose day-to-day mood is not the kind of thing a fixed file can isolate. A tool that shows per-axis results, the way Rate Cock does on public entries, lets a control shot do even more: not just "did the total move" but which axis moved, which narrows a suspected recalibration to the part of the rubric it actually touched.

How often to run it

Once per session is enough for most people - resubmitting the control every single time you submit anything else is thorough but rarely necessary, since tool drift happens on a timescale of months, not days. A reasonable cadence is once whenever you are about to make a comparison you actually care about: before judging whether a real change happened, check the control first, and only trust the real result if the control came back roughly where it usually sits.

Keep the control file itself somewhere it cannot be accidentally edited, resized or re-exported by an app that touches metadata on the way - a control shot that has silently changed format is no longer testing the tool, it is testing the file conversion, which defeats the purpose in a way that is easy to miss until the numbers stop making sense.

Read next

Full archive