Photos

The multi-angle set and its total

A score averaged across angles is smoother and less informative; the per-angle spread is where the information was.

By 4 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

A set shot from the front, the side and three-quarters averages into one number that looks tidier than any single angle's result. Tidier is the problem: averaging across angles throws away exactly the information a multi-angle set exists to capture.

What averaging does to a set of angles

Each angle reads proportion differently, because foreshortening compresses length along the line of sight by a different amount at every tilt. A front-on shot and a three-quarter shot of the same subject are not measuring the same apparent thing, and a tool that scores each and averages the results is combining two different readings into a figure that corresponds to neither.

The average also cancels information. If one angle scores noticeably higher than the others, that gap is data - it tells you which angle the subject reads best from, or which angle the tool's rubric happens to reward. Folded into a mean, that gap disappears, and the total looks like a single stable measurement of something that was never stable across the set to begin with. Machine-learning research has found the same weakness in the simple average: Shanmugam et al. (ICCV 2021), studying test-time augmentation, report that averaging predictions across transformed versions of one image "can change many correct predictions into incorrect predictions," even when overall accuracy improves.

This is a different failure from the one where a tool decides which of several photos its combined score actually reflects. That question is about which image wins; this one is about what gets lost when angles that were never comparable to each other get blended into one figure regardless of which combination rule is running.

A worked example

Say a front-on angle scores 6.8 and a three-quarter angle of the same subject, taken moments later under the same light, scores 7.6. Averaged, the set reports 7.2 - a single figure that suggests a stable middling result. Neither angle produced a 7.2. One angle read low, one read high, and the total is a number that describes neither the front view nor the three-quarter view, only their arithmetic midpoint.

A future set that happens to lean more heavily on the three-quarter angle - a slightly different tilt on the day, a photo that came out sharper - will report a higher average without the front-on reading having improved at all. The change looks like progress and is really a shift in which angle dominated the blend.

What to ask for instead

The per-angle scores, not the average. If a tool only surfaces the total, the same set can still be submitted one angle at a time, individually, to recover what the averaging step discarded.

Once you have per-angle numbers, the spread across them is worth more than the mean. A tight spread says the subject reads consistently regardless of angle, which is itself informative and rare enough to be worth knowing. A wide spread says the total was hiding a genuine disagreement between angles, and no single number - averaged or otherwise - was ever going to represent that set honestly.

This is also the more useful figure to track over time. A mean can hold steady while the underlying angles drift in opposite directions and cancel out, which is the same blind spot total scores have relative to a published breakdown - a service that reports results per axis makes this kind of cancellation visible instead of silent, and the angle case is a photo-side version of the same problem.

Where the compression actually happens

The geometry doing the damage is the same one behind why moving the camera moves the score at a single angle - foreshortening is a continuous effect of tilt, not a binary one, so every angle in a set sits on a slightly different point of the same curve before any scoring happens. Averaging treats those points as noise around a shared value; they are not noise, they are the set's actual shape. Nor are the per-angle readings smooth in the way the geometry is: Azulay and Weiss (JMLR, 2019) showed that "small translations or rescalings of the input image can drastically change the network's prediction," so even two near-identical angles can read further apart than their tilt explains.

None of this maps onto centimetres. A physical measurement is read the same way regardless of the photo's angle, because the tape does not care what a lens foreshortened - it is a different property entirely, and averaging across angles is a photo-scoring problem, not a measurement one.

A human reviewer handles a multi-angle set differently again. A person looking at several angles forms one impression but can usually say which angle informed it most, which is the exact information an averaged number destroys. That capacity to attribute a judgement to a specific angle is worth more than the smoothing an average provides, and it is why the per-angle spread, not the mean, is the number worth asking a tool to show.

Read next

Full archive