Scores

The line chart of your past scores, read carefully

A history graph mixes tool changes, photo changes and noise into one line, and the eye finds a trend in all of it.

By Updated 4 min readScores

Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how

A rating history graph shows your past scores in order, but its line can move for three different reasons - a different submission, a retuned tool, or plain noise - and the chart does not say which. A line going up looks like progress; it rarely tells that single story.

Three things are moving under one line

The submission. The most obvious source of movement, and the only one most readers assume is at work. A different photo, a different angle, a different day, and the number the tool returns changes for reasons that have nothing to do with any underlying change.

The tool. Rating models get retuned. A rubric gets adjusted, a calibration mapping gets retuned against a larger population, a new model version replaces the old one behind the scenes. None of this is usually announced, and a jump in the middle of your history graph can be the tool changing rather than anything about you or your submission changing. There is no way to tell from the graph alone which points came from which version. Hosted models do shift between versions: Chen, Zaharia and Zou (2023) found GPT-4 identified prime numbers with 84% accuracy in March 2023 and 51% in June.

Noise. Even with a fixed submission and a fixed tool, repeated ratings vary. A few tenths of movement between adjacent points on the graph is inside that range far more often than it looks like from a chart that draws a smooth line through every value as though each one were exact.

A graph that plots all your historical results as one continuous series is implicitly claiming these three sources are all one thing: change over time in whatever the tool is measuring. They are not, and the chart has no way to label which point belongs to which cause.

Why a trend needs a fixed protocol behind it

A rising or falling line only means what it appears to mean if everything except the thing you actually intended to track was held constant across the points on it. If distance, angle, light and crop drifted a little from photo to photo - which they usually do without anyone deciding to change anything - that drift is itself a source of movement in the score, indistinguishable on the graph from anything else. A trend line drawn through unprotocoled points is not evidence of a trend in the underlying subject; it is evidence that several things varied together, and the graph collapsed them into a shape that looks like one story. Keeping a series comparable over a long stretch of time is a separate discipline from reading any single graph, and a graph built without that discipline behind it cannot be rescued by careful reading afterward - the information it would need was never recorded.

Reading one honestly

Look for the shape a real trend leaves, not the shape noise leaves. Quality engineers formalise this as the control chart, whose limits, in the words of the NIST/SEMATECH e-Handbook of Statistical Methods, "are chosen so that almost all of the data points will fall within these limits as long as the process remains in-control". Noise looks like small, frequent up-and-down movement around a roughly flat line. A tool change looks like a step: several points at one level, then a jump, then several points at a new level, with the jump itself larger than the noise around either side of it. A genuine change in the submission, if the conditions were actually held fixed, would show as a gradual drift rather than a step, and it is the rarest of the three patterns to actually see in a real history graph.

If the graph shows a step, the honest question is not "what changed" but "when did the tool last update," and most services do not publish that date, which is itself worth noting about the tool rather than about you. If the graph shows scattered movement with no clear step and no clear drift, that is what noise around a stable number looks like, and the useful summary statistic is the median of the points, not the direction the last two happen to point.

What the graph cannot tell you on its own

A history graph, however smooth the line, cannot tell you whether the underlying model behind it is any good. That question belongs to the model itself, and no amount of your own retesting settles it. If the property you actually want to track changes in is length rather than a rating, a graph of rating-tool output was never going to answer that; a ruler produces a number you can trend honestly in a way a rating score cannot, because the rating scale itself is not stable underneath you. Tools that expose more than a total number make this easier to untangle - Rate Cock's per-axis breakdown lets you check whether a jump moved every axis at once, which is a tool-level step, or one axis alone, which is more likely a real change in that specific judgement. And a graph is, in the end, a poor substitute for the kind of specific, dated feedback a human reviewer gives about what actually changed between two submissions, rather than a number that leaves you guessing.

Read next

Full archive