Tools
How the written feedback differs by tool
The same score comes wrapped in very different prose depending on the tool; tone is a product decision and it changes what readers take away.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Rating tools write their result paragraphs in one of four house registers - flattering, neutral, blunt or comedic - chosen once for retention and shareability, not for your submission. That is why two tools can return the same 6.8 and leave a reader with opposite impressions.
The flattering register
The most common tone across the category, for the same reason scales themselves run generous: a user who leaves feeling good about a middling result is a user who comes back. A flattering paragraph hedges every criticism with a compliment, front-loads the positive, and buries anything corrective in a subordinate clause. It is written to be shareable, which is a separate goal from being informative and often works against it. Where a language model writes the paragraph, training can push the same way: analysing human preference data, Sharma and colleagues (2023) found that "when a response matches a user's views, it is more likely to be preferred", and that five state-of-the-art AI assistants consistently showed sycophancy.
The neutral register
Rarer, and closer to a lab-report register: the prose describes what the axes show without editorialising, states a weakness the same way it states a strength, and avoids emotional language altogether. This is the tone most useful for actually reading a result, because it does not pre-digest the number into a feeling before you have looked at it yourself. It is also the tone least likely to be screenshotted, which may be part of why it is uncommon.
The blunt register
A smaller set of tools lean the opposite direction from flattering, favouring terse or even harsh phrasing as a differentiator - the pitch being "unlike the others, we will tell you the truth." This sounds like honesty and is frequently just a different marketing angle: bluntness performs credibility the same way flattery performs kindness, and neither is evidence the underlying number is any more or less accurate than a tool with a gentler register.
The comedic register
A fourth register worth naming separately: prose written to be funny rather than either kind or harsh, leaning on exaggeration, memes or self-aware absurdity. This tone is easiest to mistake for harmless, because it does not read as flattering or as blunt, but it is doing a third thing entirely - it is optimising for the screenshot, the same audience a flattering tone chases, just through a different emotional door. A joke about a low score can travel further than a kind one, because it gives the reader something to post that is not simply "I did badly," and that shareability is the actual design goal underneath the humour.
Where the register comes from
Tone is set once, in a style guide or a prompt template, by whoever writes the copy for the product - not by anyone looking at your specific submission, and not derived from the axes in any documented way. A tool that markets itself as "brutally honest" wrote that positioning before it had ever seen your photo, the same way a tool that markets itself as "your biggest hype machine" did. Both are brand decisions made in advance, and the paragraph you get is simply that decision executed against your particular numbers.
Why the register is a separate decision from the score
The prose paragraph under a score is generated to match the number, not derived from an independent look at the submission - the axes and the total are computed first, and the wording is filled in afterward to explain them in the house style. That means the tone is chosen once, at the product level, and then applied uniformly regardless of which specific result it is describing. A tool did not decide to be gentle with you specifically; it decided to be gentle in general, and you got the house voice.
Reading around the tone
Whichever register a tool uses, the discipline is the same: read the number and the breakdown, and treat the paragraph as a caption written after the fact rather than a second, independent opinion. When the prose and the number appear to disagree, the number is the one to believe, because the paragraph was written to the number and not the other way round. The clearest demonstration of tone as a separate axis is two tools returning the identical number wrapped in prose that reads like opposite verdicts, which only makes sense once tone is understood as a house style rather than a signal.
Tone is a presentation choice, and presentation choices sit downstream of the actual computation. What the model is doing with an image has no tone at all until a separate text-generation step wraps language around the result. On the measurement side, a reported figure either states its method or it does not - there is no flattering or blunt way to write a distance, which is one advantage a tape has over a paragraph. Register matters more, not less, when the source is a person: how a human judge chooses to phrase a response is a genuine communication choice rather than a house style applied automatically, and it is worth reading as one. Some result pages keep the prose short and let the breakdown carry the weight instead - Rate Cock favours the axes over an extended paragraph, which sidesteps the tone question by leaving less prose to have a tone at all.