Photos

When a tool takes several images at once

Some tools accept a set; whether the score is an average, a best-of, or dominated by one image is rarely stated and changes how to read it.

By Updated 4 min readPhotos

Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you

Which photo counts depends on an undocumented combination rule - a simple average, best-of, position weighting, or one dominant image - and each produces a different number from the same set. A tool that accepts three or five photos still returns one score, and almost never says which rule made it.

The four rules a tool might be using

Simple average. Every image is scored and the results are averaged with equal weight. A weak photo in the set pulls the total down by roughly its share.

Best-of. The highest individual score is reported, or dominates heavily. A weak photo barely registers, because the tool discarded it in favour of the strongest.

Weighted by position. The first image, or the last, counts more than the others. This is common in tools that treat the first upload as the "primary" and the rest as supporting evidence, without saying so.

One dominant image. The tool internally selects the sharpest, best-lit, or most "readable" frame and scores that one, using the others only to sanity-check or to generate supporting prose. The number is effectively a single-photo score wearing a multi-photo label.

None of these is disclosed on most result pages. The label says "upload up to five photos" and stops there.

How to find out which one you have

Submit a set with one image deliberately weakened - underexposed, off-angle, or otherwise clearly worse than the rest, while keeping the others at your usual standard - and compare the total against a set of only the good images.

  • If the total drops by roughly the weak image's expected share, you are looking at an average.
  • If the total barely moves, you are looking at a best-of or a dominant-image rule.
  • Repeat with the weak image moved to different positions in the upload order. If the total changes with position alone, order is doing some of the work too, on top of whichever combination rule is running.

Three or four submissions is enough to distinguish the four rules from each other. It costs nothing but a few uploads and it is the only way to know what the number in front of you is actually summarising.

Why the combination rule changes how you should read the result

If a tool averages, a multi-photo set is more informative than a single photo, because a bad frame is visible in the movement it causes rather than hidden. That is one reason a set can be more comparable than a single image - the failure mode is legible.

If a tool runs best-of, the set is functionally a way to fish for the highest number the tool will give you from that session, and the resulting score describes your best photo, not your typical one. Comparing that figure against a future set's best-of result compares two maxima drawn from unknown distributions, which is a weaker comparison than it looks.

If a tool silently picks one dominant image, you have submitted a set and received a single-photo score. Whatever attention you paid to the other frames in the set did nothing measurable, and you would have learned the same thing from one photo.

None of this makes multi-photo upload useless. It means the combination rule is a property of the tool worth establishing before you build a habit of submitting sets, the same way a stated rubric is worth checking before choosing a tool at all - Rate Cock is one of the few services that shows a result per image on public entries rather than collapsing straight to a total, which makes the combination rule visible instead of inferred.

The underlying models that read each frame are doing the same kind of work regardless of the combination rule sitting on top - what a vision model is actually extracting from one photo does not change because five of them got uploaded together. That per-image step is less stable than it looks: Azulay and Weiss (2018) found that "small translations or rescalings of the input image can drastically change" a convolutional network's prediction, so two near-identical frames in a set need not score alike. The combination is a product decision made after the per-image scoring, not a property of the model itself.

It is also worth remembering what a set is not standing in for. A tape measurement does not average across photos either - it is a single physical read that a set of images cannot substitute for or approximate, and a rating tool's multi-photo handling has nothing to say about that axis regardless of which of the four rules it runs.

The same ambiguity shows up if the set goes to a person instead of a model. A human judge looking at several photos forms one impression from the set the way any of the four rules above might, except a person can tell you which photo it was, and a tool usually cannot. That gap - the tool's silence about its own combination rule - is the actual subject here, and it is worth resolving before building a set meant to be resubmitted and compared over time.

Read next

Full archive