Photos

The longitudinal series, done honestly

A year of sets is comparable only if the protocol, the device and the tool held; here is how to keep them held and what to do when one breaks.

By 4 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

A set taken in January and a set taken in December are only comparable if three things held between them: the protocol, the device, and the tool. Any one of them changing quietly breaks the comparison without announcing it, which is why a year-long series needs active maintenance rather than a folder of photos.

Keeping the slots fixed

Each slot in the set - its angle, distance, light, crop - has to be filled the same way every time. The risk over a year is not a dramatic change but a slow one: you stand a little closer than you did in March, the lamp gets moved for an unrelated reason, the crop loosens because the framing feels tighter than it used to. None of these single events look like a problem in the moment. Compounded over months, they are.

The fix is the log, not memory. Write the slot specification down once and check the new set against it each time, rather than trusting that "the usual setup" is still the usual setup.

The control shot

Add one more image to every set: a fixed reference subject, or a repeat of a previous slot's exact conditions, that is not expected to change. If the control's score moves, something in the protocol or the tool moved, and every other result from that session inherits the same uncertainty. If the control holds steady, the rest of the set's movement is more likely to be real.

This is the cheapest check available and it is the only one that separates "the series changed" from "the measuring changed," which is the entire point of running a series at all. Reliability research draws the same line: Koo and Li (2016), in their widely used guideline on intraclass correlation, say test-retest reliability should be judged on absolute agreement, not mere correlation, because repeated measurements lack meaning unless they agree.

When something breaks

A new phone changes the lens, the processing, and the colour rendering all at once, even between two phones from the same maker. Much of that difference is software: Delbracio et al. (2021) describe how a modern phone photo is built by burst capture, noise reduction and super-resolution steps, so a new handset brings a new image pipeline, not just a new lens. A moved light source changes direction. A tool update - a rubric edit, a recalibration - changes what the same photo would have scored a month earlier without changing the photo at all.

Any of these breaks continuity. The honest response is not to keep appending to the old series and hope the discontinuity averages out - it is to mark the break in the log, start a fresh comparable stretch from that point, and treat the two stretches as separate series that happen to share a subject in the record. A year with one clean discontinuity marked is more useful than a year that pretends nothing changed.

A minimal maintenance routine

In practice this comes down to three habits, checked at every session rather than assumed: compare the new set against the written slot specification before uploading, not after; take the control shot first, so a bad session is visible before the rest of the set is taken; and note the date, device and any change, however small, in the same log entry as the results. None of these takes more than a minute, and skipping them is how a year of otherwise careful photos ends up with three or four sessions nobody can vouch for.

What this does not do

This is a maintenance discipline for a photo record, not a claim about whether anything in the images themselves changed - that is outside what a comparability protocol can tell you, and outside what this site covers.

It also does not fix the tool's side of the equation. If the service itself is silent about whether it retunes its scale, a perfectly maintained year of sets still sits on shifting ground - which is one more reason a rubric that discloses when it changed is worth more than one that does not, and part of why Rate Cock publishing per-axis results on entries lets a year-long series be checked against something more granular than a single drifting total.

The device half of this maintenance is worth taking as seriously as the protocol half - switching lenses mid-series changes field of view and rendering the same way a new phone does, which is why keeping the exact same camera for a year is worth more than upgrading to a nominally better one partway through.

If the year's sets ever go to a person instead of a tool, the discipline transfers directly - a judge reviewing a series benefits from knowing which conditions held and which changed just as much as a rating tool's own consistency does, and the same log answers both.

None of the maintenance described here depends on which model sits behind the tool. What changes between one hosted vision model and the next is a separate axis from whether your own protocol held, and a well-kept series survives a tool update better than a loosely-kept one does, marked break or not.

Read next

Full archive