Archive
Everything published here
137 pieces, newest first.
-
A high total with a low axis, or the reverse
A breakdown that does not match its total is telling you about the weighting; a breakdown that contradicts itself is telling you about the noise.
-
The disagreement that actually matters
Two tools printing different numbers is normal; two tools ranking the same set in a different order is the disagreement worth investigating.
-
The longitudinal series, done honestly
A year of sets is comparable only if the protocol, the device and the tool held; here is how to keep them held and what to do when one breaks.
-
A procedure for putting two raters side by side
Same file, same day, several submissions, compare rank not number: the procedure that turns "they disagree" into something you can read.
-
A 7 that means "top 30%" and a 7 that means "7
Some tools map output onto their own past results and some onto a fixed rubric; both print out of ten and the numbers mean different things.
-
Making a submission that can go to several tools
If you intend to compare tools, the submission has to be identical to each; here is what identical means and what breaks it.
-
What a penis rater actually does
Four different mechanisms hide behind the same three names, and each one produces a different kind of number.
-
Why last year's 7 is this year's 7.5
A tool under commercial pressure has every reason to let its scale creep up and none to let it creep down.
-
What "objective" can and cannot mean here
A rating tool can be consistent, and that is the most it can be; objectivity would need a fact of the matter, and there is none.
-
Sharpness as a variable
Blur removes the detail some axes depend on; a slightly soft image scores differently from a sharp one of the same subject.
-
The examples a tool shows you before you submit
A tool's sample results are selected to sell it; they still tell you about the scale, the axes and the tone, if you read them for that and nothing else.
-
Numeric agreement with narrative disagreement
Two tools give the same 7 and describe it in opposite terms; the number is where the tools happen to overlap, the prose is where their rubrics differ.
-
Reading the gap between two retakes
Two retakes under the same conditions will not match; a small gap is normal, a large one is a signal that something in the conditions moved.
-
A single number cannot carry the weight put on it
One rating is a draw from a distribution; the useful quantity is the centre of several, and the useful question is how many is several.
-
The line chart of your past scores, read carefully
A history graph mixes tool changes, photo changes and noise into one line, and the eye finds a trend in all of it.
-
Versioning as a property of a serious tool
A rubric that changes silently makes every past score uninterpretable; a tool that dates its changes is one you can compare against.
-
Axis, defined
An axis is one component the tool scores separately before combining; the word promises independence it does not always deliver.
-
Ten is a convention, not a finding
The ten-point scale is inherited from everywhere else, and it imports assumptions the tools never checked.
-
The table people want, and why nobody can publish it
A chart mapping tool A's 7 to tool B's 8 would need both scales to be stable and both populations to be the same; neither holds.
-
Handling an outlier in a series
An outlier is either a condition slip or the tool being the tool; check the log first, and if nothing moved, keep it and use the median.
-
A test that needs no new photo
Crop one photo three ways and submit each; the spread tells you how sensitive the tool is to framing, at zero cost in retakes.
-
Same tool, two submitters, two numbers
Even on one tool, two people's scores differ by their photos, their conditions and their devices before anything else; the comparison is mostly noise.
-
Using scores to order your own submissions
A tool orders a set of your submissions more reliably than it scores any one of them; the ends of the order are trustworthy, the middle is not.
-
Removing the hand from the equation
A fixed support and a timer remove hand shake, drift in distance and drift in angle in one move; it is a repeatability upgrade, not a quality one.
-
A common claim, tested
Some tools say they normalise for angle or distance; whether they do is checkable with two photos, and the answer decides whether you can relax your protocol.
-
Working out a tool's weights from its outputs
If a tool shows a breakdown and a total, a few results are enough to tell whether the total is a plain average or something else.
-
A control for the tool, not the subject
Resubmit one fixed reference image alongside each new submission; if the control moves, the tool moved, and your new result inherits that.
-
The device question, answered
Any modern phone is enough for a repeatable submission, provided it is the same phone, the same lens and the same settings every time.
-
The advice under the score, read sceptically
Improvement tips are written to raise the next score on this tool, which is a different goal from making the next score mean more.
-
Zero variance is also information
If two genuinely different photos return the identical score, the tool may be caching, rounding coarsely or not looking closely; too little noise is a finding too.
-
The angle variable, isolated
A downward tilt foreshortens; a small change in tilt is invisible to you and not to the tool, which is why level is the only repeatable angle.
-
What a series buys and what it does not
People submit again hoping for a higher number; what more submissions actually deliver is a tighter estimate of the same number.
-
Distinguishing a difference from noise
A gap between two submissions matters when it is larger than the spread you see repeating either one; a rule of thumb, and where it fails.
-
What a paid tier actually buys
Paying usually buys more axes, more prose and a history; it does not buy a more accurate number, and sometimes it buys a kinder one.
-
The statistical reason retakes disappoint
An unusually high score is partly luck; the next one is expected to be closer to the centre, and that is not the tool changing its mind.
-
The bottom of the scale is mostly decorative
Rating tools rarely use their lowest values, and the reason is a design choice, not a compliment.
-
Where the results page works against the reader
Locked breakdowns, blurred axes, "unlock your full result": the results page is where a tool's incentives are most visible.
-
Editing as a score variable
A filter, a smoothing pass or a colour grade changes the image and therefore the number; it also makes the result incomparable with any unedited one.
-
The same rating in three wrappers
The same rubric can sit behind a website, an app or a chat bot; the format changes what you see, keep and can compare.
-
The conditions for a longitudinal comparison
Two scores months apart are comparable only if the tool, the conditions and the protocol were the same, and usually at least one was not.
-
What the written feedback in a result actually is
The prose is generated to match the number, not derived from an independent look; treat it as a caption, not a second opinion.
-
Rejection as a feature
A tool that scores everything is a tool that scores noise; the refusals tell you where its judgement starts.
-
Six questions to answer first
Same tool? Same version? Same conditions? Same protocol? More than one sample each? Only then is the comparison worth making.
-
A term from drawing, applied to submissions
Foreshortening is the compression of length along the line of sight; any tilt toward or away from the lens produces it, and the tool scores the compressed version.
-
What the decimal place in a score is actually carrying
A second digit looks like precision; in most tools it is a rounding of an internal value that was never that stable to begin with.
-
A screenshot of a result, and what it leaves out
A shared score arrives without the tool version, the conditions or the retakes; it is a number stripped of everything that gave it meaning.
-
What a tool that wants to be trusted makes visible
A described scale, a named rubric, a change log, an owner: the things a tool shows when it has nothing to hide.
-
The events that retire a reference submission
A new phone, a moved lamp, a tool update, a change of state: any of them ends the series, and the honest move is a new baseline.
-
How rubrics differ across the category
Published or hidden, three axes or twelve, real or decorative, versioned or silent: the full set of ways rubrics differ and what each choice costs the reader.
-
The tools that give you one number
A single-number tool is the easiest to read and the hardest to learn anything from; what it keeps and what it throws away.
-
The scorers that came before image tools
Before tools looked at photos, they scored questionnaires; the habit of a number out of ten with a paragraph of prose comes from there.
-
The most common confusion about submissions
A better photo scores higher; a standardised photo scores comparably. People want the first and need the second, and the tools encourage the confusion.
-
The tools that show you components
Tools with a breakdown differ in how many axes, whether the axes are real, and whether the total is derived from them; the label "multi-axis" hides all three.
-
The tools that compare instead of score
A pairwise tool never claims a value, only an order; that is a weaker claim and a more defensible one.
-
What professional capture changes and what it does not
Better gear and lighting change consistency more than they change the score; the useful question is whether the result is repeatable, and a professional setup can be.
-
Separating two sources of variance
Two tests separate them: the same file twice measures the tool; a fresh photo under the same conditions measures you plus the tool. Subtract.
-
When the totals agree and one component does not
Disagreement on a single axis is more informative than disagreement on the total; it points at a rubric difference you can name.
-
Whether the paid tier is kinder
A paid tier should show more, not score higher; if the same submission scores differently on the two tiers, that is a finding.
-
Two words used interchangeably, and why they should not be
Framing is what you include at capture; cropping is what you remove afterward. Framing changes the perspective; cropping does not.
-
Bell, skew, or pile-up: what the histogram of a tool looks like
Two tools with the same average can distribute their scores very differently, and the shape decides what a given number is worth.
-
What the common axis words usually denote
Rubric vocabulary repeats across tools; here is what each common term usually covers and how loosely.
-
A 6.1 with glowing text, or an 8 with faint praise
When the paragraph and the score point in different directions, believe the score - the paragraph was written to it.
-
Frames from video, compared with photos
A video frame is more compressed, lower resolution and softer than a photo from the same phone; it can be a series of its own, but not mixed with photos.
-
The confidence number, decoded
A confidence figure next to a score is usually about image quality or model certainty, not about whether the score is right.
-
An illustrative side-by-side, hypothetical numbers
Four hypothetical results for one submission, laid out with breakdowns, and what each disagreement turns out to be.
-
Why extreme results circulate and ordinary ones do not
The scores people show are the best of many; the ones you see are selected, and selection is why they look impossible.
-
The most useful thing a result can say
A tool that highlights what changed between two results is giving you the one piece of information a total cannot.
-
Why visible results change what a reader can check
A tool that lets you see other submissions' breakdowns lets you check its axes, its spread and its consistency; a tool that hides all results asks for faith.
-
Whether the breakdown is real
Some breakdowns are computed per axis and some are one number split into six for show; there is a way to tell from the outside.
-
The same result presented three ways
A 7.4, a silver badge and a paragraph of praise can be one output; each presentation makes the reader believe something different.
-
The population a score is implicitly compared to
Any score that means "better than most" depends on who "most" is, and tools rarely say.
-
When two tools agree on the digit and nothing else
Two sevens from two tools are not the same seven; the agreement is on the display, not on the claim.
-
The mean of two incompatible numbers
Averaging a 6.5 and an 8 from two tools produces a number on no scale at all.
-
How the written feedback differs by tool
The same score comes wrapped in very different prose depending on the tool; tone is a product decision and it changes what readers take away.
-
The top of the scale, examined
A perfect score is a claim that nothing could move the number upward; almost no tool is built to make that claim.
-
Choosing the unit of comparison
A single photo is easier to reproduce; a set is more informative and harder to hold constant. Which one to standardise on depends on what you want to compare.
-
Image size as a variable
Below a threshold, resolution costs detail and moves the score; above it, more pixels change nothing except the upload time.
-
When agreement is not independence
Many tools sit on similar hosted models; two of them agreeing may mean one judgement seen twice.
-
The grade analogy, and why it misleads
Readers import school-grade meaning into rating scores; the two scales share digits and nothing else.
-
Two properties people conflate
A tool that gives the same photo the same score every time is consistent; whether the score is right is a separate question with no clean answer.
-
The adjective attached to the number
A word next to a score is a threshold decision; where "good" starts is the tool's call, and it is usually generous.
-
A survey of uncertainty presentation across the category
Some tools show a confidence value, some a range, most nothing; the choice tells you how the tool wants its number read.
-
Where the gap between raters is widest
Two tools tend to agree in the middle and diverge at the ends, because the ends are where scale design differs most.
-
Composite, defined
A composite is a total built from parts by a rule; the rule is the part that matters and the part you are least likely to be shown.
-
The crop as a score variable
A crop decides what fraction of the frame the subject fills and what context surrounds it; both are inputs, and both move the number.
-
The systematic tilt in rating tools, and where it comes from
A tool that scores low loses users; a tool that scores high keeps them. The centre of the scale moves accordingly.
-
The tools that never look at a photo
Some raters score a questionnaire, not an image; the number is a function of self-report, and it belongs to a different category entirely.
-
What changing your distance to the lens does to a result
Move the camera and the proportions the tool reads change with it; distance is the single largest source of score movement that has nothing to do with the subject.
-
Where a score gets rounded, and what that hides
A rounded score can move a whole point on a change too small to see; rounding is a threshold, and thresholds create false jumps.
-
The result designed to be posted
A share card keeps the number and the brand and drops the breakdown, the conditions and the caveats; it is a result optimised for a different reader.
-
When a tool takes several images at once
Some tools accept a set; whether the score is an average, a best-of, or dominated by one image is rarely stated and changes how to read it.
-
A total and a breakdown are not the same information
A single number compresses several judgements and throws away which one moved; a breakdown keeps that, at the cost of being harder to brag about.
-
Drawing the boundary of the category
A tool that returns a number is a rater; a tool that returns an edited image or a compliment is something else wearing the name.
-
A small test for multi-photo tools
Some tools weight the first image or the last; reorder a set and resubmit, and if the score moves, order is a variable you must fix.
-
The comparative labels, decoded
A tool that says you are above average is comparing you to something; which something changes whether the label is worth anything.
-
Where a tool's standard comes from
A rubric encodes someone's decisions about what counts; those decisions are the standard, and no tool found them in nature.
-
What an old file tells you when you send it again
Sending last year's file today measures how the tool has changed; the photo did not, so any difference is the tool's.
-
A short list to run through
What scale, what rubric, what population, what version, how repeatable, who owns it. Six questions and why each matters.
-
The processing screen, and what it claims
Analysing 27 features..." is copy, not a progress report; what the waiting screen says and what is actually happening are unrelated.
-
Frame orientation as a variable
Orientation changes fill ratio and how the tool crops on input; it is a small variable, and small variables are why protocols exist.
-
The tools that give you a place, not a number
A rank-only tool tells you where you fall among its users and nothing about the scale; it is honest in one way and opaque in another.
-
The midpoint of the scale is not the middle of the results
Most tools centre their output somewhere between six and seven, and reading a six as below average is the most common misreading there is.
-
The axis count as a product choice
More axes look more rigorous and are harder to keep independent; the count tools chose tells you what they were optimising.
-
A set with different lighting, angles or states in it
A set whose images were taken under different conditions has no single condition; its score describes nothing you can reproduce.
-
What the noun in "your score" refers to
A rating is a property of an image under conditions; the pronoun in "your 7.4" is doing a lot of quiet work.
-
Two lens terms, defined
Field of view is how much the lens takes in; crop factor is how that compares to a reference; both change how proportion is rendered at a given distance.
-
What happens when many submissions crowd the top
If a tool scores generously, its top two points have to hold most of the population, and differences up there stop meaning much.
-
The one submission everything else is compared to
A baseline is a submission taken under recorded conditions that you can repeat; without one, no later result has anything to be compared with.
-
What the tool says it is scoring
Tools name the thing they score differently, and the name sets what a reader thinks the number is about; usually the number is the same underneath.
-
Why "proportion" on one tool is not "proportion" on another
Two tools using the same axis name are not scoring the same thing; the name is a label on a scale each tool built itself.
-
Where the rating tool category came from
The ten-point scale, the share card and the crowded middle all predate the current tools; the category inherited its shape before it inherited its models.
-
Why a 6 on one axis and a 6 on another may not match
Axes are printed on the same ten-point scale and often centred differently; a six on a generous axis is a different claim from a six on a strict one.
-
Why a note next to each submission matters
A score without its conditions cannot be compared later; the log is what turns a pile of results into a series.
-
Why the capture path changes the image
A mirror adds a reflection layer and a distance; a front camera is a different lens with different processing. Neither is wrong, but they are not interchangeable.
-
The most common score in the category, and why
Seven is where generosity, a crowded middle and a rounded display all land at once.
-
Getting from six numbers to one
Equal weights, fixed weights, and non-linear combinations all produce a total out of ten; each one makes a different submission look best.
-
The correlation failure in multi-axis tools
If all six axes rise and fall together across your submissions, the tool has one judgement and six labels.
-
Why two tools disagree on the same photo
Submit one image to two raters and expect two answers. The gap is mostly calibration, and calibration is not accuracy.
-
The label as a tool-category claim
Nearly every rating tool now says AI; the label tells you almost nothing about how the tool differs from the next one that says it.
-
The term this site keeps using
A standardised submission is one taken under recorded, repeatable conditions, chosen for comparability rather than for score.
-
The countdown, the spinner, the drum roll
The delay before the score is presentation, not computation, and it raises the stakes of a number that does not deserve them.
-
How the tool makes money and what that does to the score
A subscription tool wants you back, a credit tool wants you rescoring, an ad tool wants you sharing; each shapes the number a little differently.
-
The case for a short series every session
A single submission is one draw; five under identical conditions give you a centre and a spread, and the spread is the part that keeps you honest.
-
What a score out of ten is made of
The scale is the most familiar part of any rating tool and the least examined. It is not a measurement, and it does not behave like one.
-
Agreement is weaker evidence than it looks
Two tools converging on the same score is reassuring and mostly uninformative if they were built the same way and given the same photo.
-
Every presentation choice on a result page, and what it does
Number, badge, colour, label, bar, prose, animation, share card - the full set of presentation choices and how each one steers the reader.
-
How much weight a rating tool result deserves
The number is real, the confidence is borrowed. This is the full argument for how much to trust a rating and where that trust runs out.
-
A complete protocol for comparable submissions
Fix the device, lens, distance, angle, light, crop, background and state; record it; take several. The whole protocol, and why each line is there.
-
Every part of a rating result, and what each one is for
A result page has four or five distinct components, each produced differently and each deserving a different level of trust.
-
Everything about submitting more than one image
Sets, slots, combination rules, order, mixed conditions and series over time: the complete guide to making multi-image submissions comparable.
-
A taxonomy of rating tools by what they output
Single-number, multi-axis, pairwise, rank-only and quiz-style: five shapes, five different claims, and five different ways to be misread.
-
What can and cannot be compared, and how
Across tools, across time, across people, across photos, the complete map of which comparisons hold and what each needs to hold.
-
A method for judging any rating tool before trusting it
Rubric, calibration, consistency, presentation, disclosure: five checks a reader can run on any tool without special access.
-
How to find out how much of your score is you
Identical file, retake, recrop, control shot, short series: the complete method for measuring how repeatable your results are and where the noise comes from.
-
Standardising a submission, not flattering it
There is a difference between making a photo score higher and making a score mean something. Only one of them survives a second attempt.
-
Reading a rating honestly
The number arrives with a confidence it has not earned. Six rules that put it back where it belongs.
-
The five things that actually separate rating tools
The model underneath is rarely the difference. Everything above it is.