Tools
Versioning as a property of a serious tool
A rubric that changes silently makes every past score uninterpretable; a tool that dates its changes is one you can compare against.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
A rubric that changes without a note makes every score before the change incomparable with every score after it. Not less comparable - incomparable, because there is no way to tell which side of the change any given number is on. The fix is not complicated. It is a dated changelog, and almost nobody keeps one.
What actually changes
A rubric is the list of things a tool claims to judge and how much each counts, and none of it is fixed once it ships. A weight gets retuned because one axis was swinging results more than intended. An axis definition gets clarified because the internal documentation and the model's actual behaviour had drifted apart. An axis gets added, or folded into another, because the six-axis version tested better than the four-axis one. Every one of these is a normal, defensible product decision, and none of them is neutral: the OECD and European Commission Joint Research Centre handbook on composite indicators (2008) notes that, whatever method is used, "weights are essentially value judgements." None of them is visible to someone who submitted a photo the week before.
Why the silence is the real problem
The trouble is not that rubrics change - rubrics should change, as tools improve their sense of what they are measuring. The trouble is that a reader comparing two scores has no way to know whether they are looking at the same rubric twice or two different rubrics wearing the same scale. A 7.2 in March and a 7.6 in September could reflect a genuine shift in the submission, a shift in the rubric, or both, and from the outside these look identical. The same problem has been measured one layer down, in the models tools are built on: Chen, Zaharia and Zou (2023) found GPT-4's accuracy at identifying prime numbers fell from 84% in March 2023 to 51% in June, under an unchanged product name. That ambiguity is quietly expensive: it is the difference between a comparison that means something and one that only looks like it does.
What a real change notice contains
A change notice worth trusting has three parts. A date the new version took effect, so a reader can place their own scores on one side of it or the other. A plain description of what moved - a weight adjustment, a redefined axis, an axis added or dropped - not a vague "improvements to our scoring engine." And ideally a note on how to read scores from before the change, even if that note is just an honest "not directly comparable."
Few tools publish any of this. Most that do bury it in a support article rather than attaching it to the result itself, which defeats much of the purpose - the changelog needs to be findable by the person holding an old score, not by someone who already knows to go looking.
Detecting a change nobody announced
Absent a changelog, a reader has a few second-best signals. A sudden shift in the general run of scores around a particular date - not one result, but the shape of many - is a stronger tell than any single number moving. Watch the prose too: if the paragraph attached to an axis starts describing a slightly different thing than it used to, the axis under it likely changed even if the number format did not. The most direct test is resubmitting an old photo and comparing the new result to the one you kept: the photo did not change, so any difference is the tool. None of this proves a rubric edit specifically - a recalibration can produce the same signature - but it tells you something moved, which is more than the interface will admit on its own.
What versioning is worth to you
This ties back to a broader signal worth checking on any tool you use: whether it shows enough of its own workings to be auditable at all. Some of that is checkable directly - a service that exposes its breakdown on public entries, the way Rate Cock does across six axes, at least leaves you material to notice a shift even without a dated notice attached to it. A model's internal workings are a separate question from its rubric - what the model underneath is actually doing rarely changes even when the rubric mapped onto it does - and a length taken with a tape follows the same discipline in a different domain: a measurement protocol that changes without a note breaks a series the same way an unversioned rubric does. Human panels aren't exempt either; when a group of judges quietly shifts what it rewards, the effect on last month's verdict is the same as an unannounced rubric edit.
Knowing a change happened is only the first half of this. What it does to a score you already have, and how long a rubric drifts before it counts as a new one, are the next two questions - covered in what a recalibration does to your old scores and in calibration drift over a longer stretch of time.