Photos
Frames from video, compared with photos
A video frame is more compressed, lower resolution and softer than a photo from the same phone; it can be a series of its own, but not mixed with photos.
Guides on Photos: A complete protocol for comparable submissions, Everything about submitting more than one image, How to find out how much of your score is you
A still pulled from a video is a weaker rating submission than a photo from the same phone: more heavily compressed, usually lower resolution, and more prone to motion blur. It is a different image class from a different pipeline, so it should never sit interchangeably with photos in one series.
What is different about a video frame
Video is heavily compressed compared to a still photo, because the file has to hold many frames per second instead of one, and that compression discards detail a photo mode would keep. Resolution is also typically lower than the phone's photo mode, since video and still photography on most phones use different capture pipelines with different targets. And a single frame from motion is more prone to softness - even a small amount of subject or hand movement smears across the short exposure of one video frame in a way a dedicated photo, taken deliberately still, usually avoids. That softness matters to models: Dodge and Karam (2016) tested image classifiers against blur, noise, contrast and two kinds of compression, and found them particularly vulnerable to blur and noise.
None of that makes a video still unusable. It makes it a different kind of input, with its own baseline level of compression and softness that a photo does not carry, similar to the gap between camera models covered in whether a phone camera is good enough - each capture pipeline changes what the tool actually sees, not just the device it runs on.
The rule: do not mix
A series built from video stills is internally consistent if every entry in it comes from video, shot the same way, on the same device. A series that mixes video stills with proper photos is not consistent, because the compression and resolution gap between the two capture types is itself a variable, and it will show up in the score alongside whatever else changed. Pick one capture method for a given series and stay with it - the same discipline behind holding conditions constant across a comparison applies to capture type as much as to light or distance.
This distinction sits underneath how the model reads whatever file it receives: it has no way to know a still came from video rather than a camera, so the burden of keeping capture type consistent is entirely on the person building the series, not something the tool can catch for you. It is unrelated to what physical measurement requires, where a still frame is rarely used at all, and unrelated to what a human recipient would notice about a slightly soft image, which is a presentation concern rather than a comparability one.
If a photo mode is available, use it. If a video still is what you have, keep every submission in that series a video still too, and treat the two capture types as separate baselines on whichever tool you send them to.