Photos

A small test for multi-photo tools

Some tools weight the first image or the last; reorder a set and resubmit, and if the score moves, order is a variable you must fix.

By Updated 2 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

Photo order in a set can matter: some tools treat the first or last upload differently, so the same images in a different order can return a different total. Most result pages give no indication of this, which is why it is worth one test rather than an assumption.

Why order could matter at all

A tool that treats the first upload as the "primary" image and the rest as supporting context is making an order-dependent choice without saying so. So is a tool that gives the last upload extra weight because it assumes a set is arranged worst-to-best, or best-to-worst, by convention. Neither assumption is stated on the upload screen, and both are plausible enough implementation choices that it is worth checking rather than assuming.

This sits next to, but is separate from, the question of which image a multi-photo score actually reflects. That question is about combination rule - average, best-of, dominant image. This one is about whether position within the set changes the outcome on top of whatever combination rule is running.

The test

Submit the same set twice, identical images, different order. If the total is unchanged, order is not a variable for that tool. If it moves, order is doing work, and a comparable set now needs a fixed upload order recorded alongside its other conditions - the same discipline as defining each slot in a set before it gets reused.

This costs one extra submission and settles the question completely, which is a better use of effort than guessing.

What to do if order matters

Pick an order - front first, say, or angle order consistently - and keep it every time you submit a set. Treat it exactly like distance, angle or light: a condition to log, not a detail to leave to whichever image happened to be selected first on your phone.

A tool that shows per-axis results on its entries, as Rate Cock does, at least gives you something to check an order-dependent result against, since a genuine order effect should show up consistently across repeated tests rather than as a one-off. An order effect can come from the model as well as the product layer. A tool that scores each image separately is order-blind at the model level, but one that sends the whole set to a multimodal model in a single prompt is not - how these models read their input matters here, and Zheng et al. (2023) document position bias in language models used as judges.

None of this has an analogue in a physical read: a tape measurement has no upload order to test. People are not immune either. Bruine de Bruin (2005) found that in the Eurovision Song Contest ratings increased with serial position, and replicated the effect in World and European Figure Skating judging - a reminder that a judge's read of a submission has its own order effects, just different ones from an automated combination rule.

Read next

Full archive