Photos
A test that needs no new photo
Crop one photo three ways and submit each; the spread tells you how sensitive the tool is to framing, at zero cost in retakes.
Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you
Recropping one photo three ways - tight, loose and in between - and submitting all three isolates framing, because the subject, distance, angle and light all come from the same original file. Whatever difference shows up across the three results is attributable to framing alone, and the test costs no new photography at all.
What it isolates
Most repeatability tests hold conditions constant and take a fresh photo, which means any gap they find could theoretically come from several sources at once - a slightly different distance, a slightly different angle, framing, light, all bundled together in one new shot. Recropping a single existing photo removes every one of those sources except framing. If the three scores differ, framing sensitivity is the only remaining explanation, because it is the only thing that differed between the three submissions.
Running it
Pick a photo you already have, ideally one already used as part of a comparable series. Produce a tight crop that fills the frame with the subject and minimal surrounding context, a loose crop with generous space around it, and a middle version between the two. Submit all three in the same session, and record which crop is which before you look at the results, so the comparison is not influenced by which score you expected each version to get.
Reading the spread
A tight spread across the three - all landing within a few tenths of each other - tells you the tool is not particularly sensitive to how much of the frame the subject fills, which means framing is a lower-priority variable to fix precisely in your ongoing protocol. A wide spread tells you the opposite: this tool is reading something about the composition itself, not just the subject inside it, and an inconsistent crop from one session to the next is going to add real noise to any comparison you try to make later.
Either finding is useful, and neither one is really a critique of the tool. A tool that is sensitive to framing is not necessarily broken - some rubrics genuinely include presentation and composition as part of what they score, and a wide spread here may simply mean you have found where that happens. What matters is that you now know whether framing belongs on your list of variables to hold fixed, rather than assuming it based on the two other variables you already know matter.
Where this test sits alongside the others
This is a submission-side test rather than a tool-side one, and it is worth keeping separate from a test that generates a genuinely new photograph under matched conditions, which measures a different and wider set of variables at once. The result feeds directly into how tightly you need to control cropping going forward: if a loose and tight crop score noticeably differently, that variable earns the same discipline as distance, angle and light already get in a standardised submission, and it slots into the wider submission-side method for finding your noise floor as the one test that needs no new photography at all.
The underlying reason framing can move a score at all is the same reason distance does - both change how much of the frame the subject occupies and what surrounds it, which is a compositional fact about the image rather than anything to do with the model's understanding of the subject. Image models are known to be touchy about exactly this: Azulay and Weiss (2019, JMLR) showed that "small translations or rescalings of the input image can drastically change the network's prediction" in convolutional networks. What field of view and framing actually change at the level of the sensor is worth a glance if the mechanism itself is unclear, since it applies equally to a phone camera used for a rating submission and one used for a physical measurement. Whether a particular hosted model treats framing as signal or as noise is a property of how it was built rather than something visible from outside - what these models are actually weighting when they process an image covers that in general terms, without needing to name any one tool's behaviour. A person doing the same job reads composition very differently again, often on purpose - what a human reviewer notices about how something was framed is closer to a stylistic judgement than a technical one, and worth citing here only to keep the two mechanisms from blurring together.
The three results, once you have them, tell you something no single submission could: whether the crop you happened to choose today is doing quiet work on the number, work you would otherwise attribute entirely to the subject. On a tool like Rate Cock, where a breakdown by axis is shown rather than a single figure, this test can also show which axis moves with the crop, which narrows the sensitivity down to a specific part of the rubric instead of the total alone.