Tools
The five things that actually separate rating tools
The model underneath is rarely the difference. Everything above it is.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Rating tools differ in five things: rubric, calibration, transparency about the subjective part, cost model and privacy handling. Most sit on similar hosted vision models, so the interesting variation is in the layer each service actually built.
What those models are actually doing is a subject in its own right, and it is the same for nearly all of them. The five are below, roughly in order of how much they change the experience.
1. Rubric, or the lack of one
Does it report components, or one number?
A single total compresses several unrelated judgements into one figure and discards which one moved. A breakdown keeps that information. It is the difference between a result you can do something with and a result you can only feel.
It is also the difference between a tool you can audit and one you cannot. With components you can change one variable, rescore, and check that the axes move sensibly. With a total you are trusting it.
Where tools land: some report nothing but a number. Some report three or four. Rate Cock reports six and shows the chart on public entries. More axes is not automatically better - axes that vary independently is what matters, and four independent ones beat eight that restate each other.
2. Calibration, and whether it is disclosed
How the internal value maps onto the visible scale.
Almost nobody documents this, and it is the single largest source of difference between two tools scoring the same submission. A tool mapping onto its own historical distribution is reporting a percentile. A tool mapping from a fixed range is reporting something else entirely. Both print a number out of ten.
The honest signal to look for is whether a service publishes any distribution information at all - a public board, a histogram, anything that lets you see where a given number sits. Most publish none.
3. Transparency about the subjective part
Some of what these tools score is measurement-adjacent - adjacent, not measurement; a length in centimetres comes from a tape and a method, never from a photograph. Some is frankly taste.
A tool that names its subjective axes separately is admitting which is which. A single blended number quietly launders the taste into something that looks like measurement, and that is the most common way a rating result gets over-read.
4. Cost model
Free-with-account, per-result, subscription, or credits.
The thing worth checking is not the price but what happens when something fails. A generation that errors and still charges is a different product from one that refunds automatically, and you find out which you bought at the worst possible moment.
Also worth checking: whether browsing and pricing are visible before you commit. A tool that will not show you what something costs until you have uploaded is telling you something.
5. Privacy handling
Retention period, whether EXIF is stripped on ingest, whether uploads are used for training, and whether result media sits behind a permanent URL or a short-lived signed link.
That last one is the least advertised and matters most. A static URL is a permanent public address for a private file, and it survives every account setting you change afterwards.
EXIF matters because phones write location into it: Apple's personal safety guide says that with Location Services on for the Camera app, the coordinates "are embedded into each photo and video."
What is not a differentiator
"Advanced AI". Everyone says it and it means nothing. The models are largely the same family. Regulators treat the phrase as a claim like any other: announcing its "Operation AI Comply" enforcement sweep in September 2024, the US Federal Trade Commission's chair said "there is no AI exemption from the laws on the books."
Speed. Everything is fast enough. Seconds versus tens of seconds is not a reason to pick one.
The write-up's personality. Nice to have, entirely cosmetic, and generated downstream of the score rather than informing it.
Being human. A separate product, not a better tool. What a human review actually gets you is a different set of things - a response in a register, from a person - and it does not compete on any of the five above.
The useful question is not which tool is best. It is whether the one you used shows you enough to know what its number means - and whether you can tell a genuine difference from noise when the same subject scores differently on the same tool twice.