Scores
The conditions for a longitudinal comparison
Two scores months apart are comparable only if the tool, the conditions and the protocol were the same, and usually at least one was not.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
A March score and a June score are comparable only if the tool version, the submission conditions and the protocol all stayed fixed between them. Usually at least one did not, and then the gap measures that change rather than anything in the photos.
The tool and its version
The rubric behind a tool is not a fixed object. It gets edited, retuned, and rescored against a shifting population of submissions, often without a visible changelog. When a tool recalibrates, a March score and a June score can differ for reasons that have nothing to do with what changed between the two photos - the scale itself moved under both numbers. Silent change is documented even in the largest hosted models: Chen, Zaharia and Zou (2023) measured GPT-4 identifying prime numbers with 84% accuracy in its March 2023 version and 51% in June, and concluded that the "same" service can change substantially in a short time. If the tool does not date its rubric changes, you cannot rule this out, and the honest move is to treat any gap larger than expected as ambiguous rather than as evidence of anything.
The submission conditions
Distance, angle, light, crop, background - if any of these drifted between the two attempts, the comparison already has a second variable in it before the scale has a chance to matter. What has to be fixed for a submission to be standardised in the first place is the same list that has to hold across a gap of months, and months are exactly the timescale over which a setup quietly changes: a different room, a different season's light, a new phone. Camera and lens geometry is one of the largest single sources of that drift, because a phone upgrade changes focal length and processing at once, invisibly, on a schedule that has nothing to do with your protocol.
The protocol itself
Beyond the physical setup, the procedure has to match: same number of submissions, same handling of outliers, same choice of which result you write down as "the" score for that period. A single best-of-three read as one point and a single first-try read as another point are not measuring the same thing, even on an unchanged tool with an unchanged setup.
Sample size across time
A single score in March against a single score in June compounds two sources of noise into one gap: whatever the tool's ordinary repeat variance is, doubled across two separate single draws taken months apart. A small series at each end point - three or four submissions under the recorded protocol, rather than one - narrows both estimates and makes the eventual gap easier to trust. This costs a few extra minutes at each check-in and is the single cheapest thing to add to a longitudinal comparison that is otherwise sound. It is also the step people skip first, because a single number feels sufficient in the moment and the cost of that shortcut only shows up later, when the comparison it was meant to support turns out to rest on two lone draws instead of two estimates.
What a dead comparison looks like
If the tool changed version, the setup changed, or the protocol changed, the comparison is not weakened - it is broken, and the honest response is to say so rather than report a gap as if it still meant something. This is a stricter standard than most people apply, because the two numbers still print on the same page and invite comparison by proximity alone. A checklist run before any two scores are compared makes this concrete: tool, version, conditions, protocol, sample size. Fail one and the comparison is not degraded, it is void.
Where to log this so it does not become a guess
The fix is not vigilance, it is a written record: what tool, what version if visible, what distance and light, taken on what date. A physical measurement kept the same way is the model - method and result belong on the same line, or the result is not data six months later. Rate Cock's own what your score means explains a single result on that platform; a log across visits is what turns a series of those single results into something you can actually compare, which is a different job than reading any one of them. None of this is about whether anything about the subject changed - that question is outside this comparison entirely, and a human reviewer working the same question over time faces the identical requirement to hold the setup still before reading anything into a difference.
Two numbers months apart are comparable exactly as far as what produced them stayed the same. Check that first. The gap means nothing until you have.