Photos
What an old file tells you when you send it again
Sending last year's file today measures how the tool has changed; the photo did not, so any difference is the tool's.
Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you
Resubmitting an old, unedited photo to the same tool tests the tool itself: the file is the same bytes as before, so any difference between the old result and the new one has to come from the tool changing, not the subject or the photo.
Find the original file you submitted months or a year ago and send it through again today. Sending in that old file is a different act from sending in a new one of the same subject, and keeping the two apart in your own record is what makes this test readable at all.
What it isolates
This test removes every variable a fresh retake would carry with it - a slightly different angle, a slightly different distance, light that was never going to match exactly a year apart. The retake test measures your process alongside the tool; resubmitting an old file measures the tool alone, over a stretch of time long enough for the tool itself to have plausibly changed. Nothing about the file gives the tool an opportunity to see anything new, which is exactly what makes any observed difference so specific.
Why it needs time to be useful
Run this test a day after the original submission and you are mostly re-running the identical-file reproducibility check, which measures short-term noise rather than anything resembling drift. The value in this particular version comes from letting real time pass - months, ideally closer to a year - during which a service has had the opportunity to retune its underlying model, adjust its rubric, or shift its rounding and display logic. A gap that shows up over that kind of interval is a different finding from a gap that shows up between two submissions five minutes apart, even though the mechanics of running the test look identical.
Reading the result
A score that lands close to where it landed originally is reassuring in a specific, narrow way: it tells you the tool, at least for this one image, has not moved much over the interval you tested. It does not tell you the tool has never changed anything - a rubric edit that happens to leave this particular photo's result roughly where it was is still a rubric edit, just one that was invisible on this one file.
A score that has shifted noticeably is more informative, and worth taking seriously precisely because the photo gave the tool nothing new to react to. The direction of the shift is worth noting too - a broad upward creep across old files resubmitted this way looks different from one file moving up and another moving down, and the pattern across several old files tells you more than any single one does.
How this complements a standing control
This test works well as an occasional check run against files you already have lying around, which makes it cheap and retrospective - you do not need to have planned for it in advance. A photo set aside specifically to be resubmitted on a schedule is a related but more deliberate version of the same idea, and it is worth treating as its own procedure with its own logic rather than folding the two together here. What both share is the underlying principle: a photo that provably has not changed is the only kind of input that can isolate tool-side movement this cleanly, whether you built that input on purpose or found it in your own submission history.
Any old file that returns a genuinely different result also raises a separate question worth naming rather than answering here - what happens to a stored score once the tool it came from has recalibrated is a consequence of exactly this kind of drift, and is worth its own look before you treat the old and new numbers as comparable at all.
The tool side of this drift
None of this test explains why a tool changed, only that it did. What actually happens inside a hosted model that causes its output to shift over a long interval, without the rubric itself being edited is a mechanism worth citing rather than covering here, since it belongs to the model's own behaviour rather than to anything the submitter did. That it happens is documented: Chen, Zaharia and Zou (2023) found that "the behavior of the 'same' LLM service can change substantially in a relatively short amount of time", with GPT-4's accuracy at identifying prime numbers falling from 84% to 51% between its March and June 2023 versions. A parallel form of this test exists for a physical measurement too, in a much simpler form - a method that has not changed should return the same reading on the same body a year apart, and if it does not, the method or the tool measuring it is what moved, not the subject. A human reviewer resists this kind of test almost entirely, since a person asked to look at the same photo twice, a year apart, brings a memory and a mood the second time that a model does not - what actually varies in a person's read of returning material is a different problem from tool drift, and worth keeping distinct from it.
An old file, resent unchanged, is one of the cheapest tests available for catching movement you would otherwise only notice by feel. On Rate Cock, where a per-axis breakdown is shown rather than a total alone, running this test can also show which axis moved rather than only the overall number, which narrows down what kind of change happened on the tool's side rather than leaving it as a single unexplained shift.