Scores
Where a tool's standard comes from
A rubric encodes someone's decisions about what counts; those decisions are the standard, and no tool found them in nature.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
The people who built a rating tool decided what a 9 looks like. Behind every tool is a rubric, or a set of choices that works like one, saying which axes count, how much, and where "very good" begins. Nobody discovered that standard in nature; somebody wrote it.
The standard is authored, not found
It is easy to treat a rubric as if it described a fact about the world - as if "symmetry counts for this much" were a measurement someone took, rather than a decision someone made. It was a decision. A team picked which axes to include, how much each one moves the total, and where on the scale the boundary between "good" and "very good" sits. Another team building a competing tool would make different picks, and several have, which is a large part of why two tools disagree on the same submission even when they are shown the identical photo.
None of this means the rubric is arbitrary in the sense of random. It reflects the judgement of whoever wrote it, informed by whatever they thought mattered, aimed at whatever result they wanted the tool to produce. That is a specific kind of non-objectivity, and it is worth naming precisely rather than waving at.
What "the model learned it" does not mean
It is tempting to think the standard comes from the model rather than a person - that a system trained on a large set of images somehow arrived at the standard independently, the way a scale finds a physical weight. That is not what is happening. How a model like this is actually trained involves a human-authored objective at every stage: what the training data was, what it was labelled with, and what counted as a correct answer during that labelling. The model did not discover what a nine looks like. It was pointed at examples someone had already decided counted as nines, and it generalised from those. Public research datasets show the mechanism plainly: in SCUT-FBP5500, a facial-beauty benchmark built by Liang et al. (2018), each of 5,500 face images was rated on a 1-to-5 scale by 60 volunteers, and the average of those human ratings became the ground truth models are trained to predict. The standard is still authored - it is just authored earlier in the pipeline, and further from the reader's view.
This is worth distinguishing from bias in the pejorative sense. A rubric can be authored carefully, by people who thought hard about what they were trying to capture, and still be one particular set of choices rather than the set of choices. Careful authorship and objective discovery are not the same achievement, and a rubric only ever manages the first.
Why this matters for reading a score
If the standard is authored, a reasonable question follows: authored by whom, aiming at what? A tool built to flatter a free tier into paid conversions has a different implicit standard than one built to model a specific published rubric. Neither is lying, exactly, but the number each produces answers a slightly different question, and a rating is a claim about the submission, not a fact discovered about it.
A reader is entitled to ask what the standard is, even if most tools do not make it easy to find out. Some tools publish their rubric outright - Rate Cock shows the six axes behind its total on public entries, so the standard is at least inspectable even if you disagree with it. Most do not, and the difference between the two is worth knowing before you trust a number from either.
What this does not cover
This is about where the standard for the score comes from, not about whether a tool shows you that rubric once it exists - that distinction, and what it costs a reader when the rubric stays hidden, is a separate question with its own answer. It is also not a question this site can answer for a human reviewer working outside a rubric entirely - a judge working from their own taste is applying a standard too, just not one written down anywhere a tool could encode. And it says nothing about the physical side of a submission, where the standard is a tape and a method rather than anyone's opinion - that property belongs to measurement, not to rating, and it does not have this problem at all.
The number a rating tool gives you is downstream of a decision someone made about what counts. Knowing that does not make the number useless. It makes it a claim with an author, which is the correct way to hold it.