Tools
Trends in the category, discounted appropriately
More axes, more prose, more confidence displays, more history; which of these would make results mean more and which are decoration.
Guides on Tools: Every presentation choice on a result page, and what it does, A taxonomy of rating tools by what they output, A method for judging any rating tool before trusting it
Of the trends visible in rating tools in 2026, only three would make a score mean more: independent axes, visible history, and confidence displays that measure score uncertainty. More prose, reveal animations, streaks and friend rankings mostly improve engagement. AI tips and cross-tool comparisons could go either way.
Would improve comparability
More axes, if they are independent. A tool moving from a total to a breakdown, or from three axes to six, improves what a reader can check - but only if the added axes actually vary separately. A breakdown where every axis moves together is one judgement wearing several labels, and adding axes without checking for that just adds decoration with extra steps.
History and retakes, kept visible. A tool that stores and shows your past results, rather than showing only the latest one, makes the kind of repeatability checking this site recommends possible without you maintaining your own log. This is a real gain, conditional on the history staying accurate across any recalibration the tool does - a stored old score describing a scale that no longer exists is worse than no history at all if the tool does not flag it. Model changes are not rare events either: Chen, Zaharia and Zou (2023) measured GPT-4 answering the same prime-number questions with 84% accuracy in March 2023 and 51% in June, and concluded that model behaviour can shift substantially within months.
Confidence indicators, if they measure the right thing. A displayed confidence value is only useful if it is telling you about score uncertainty rather than image quality or model certainty on an unrelated dimension. That bar is higher than it looks, because a model's raw confidence is not trustworthy by default: Guo et al. (2017) found that "modern neural networks, unlike those from a decade ago, are poorly calibrated". Most current confidence displays measure the latter, which limits how much this trend is actually delivering yet, even where it appears to be moving in a useful direction.
Would mostly improve engagement
More prose. Longer generated write-ups make a results page feel more substantial without changing what it tells you, because the prose is generated to match a number that was already fixed before the paragraph was written. This is the same convention inherited from well before current models existed, extended rather than reconsidered.
Streaks and reveal animations. A delay before the number appears, or a counter that rewards repeated submissions, changes how a result feels to receive without changing anything about what it means. These are retention mechanics borrowed from adjacent product categories, not improvements to the rating itself.
Social and comparison features that hide the underlying number. Features that surface a rank among friends or a percentile against a curated group can obscure rather than clarify, particularly when the reference population behind the comparison is never named.
Two trends that could go either way
AI-generated improvement tips. A tool that tells you what to change to raise your next score is either giving you something to test against a real axis, or it is giving you generic copy that would apply to almost any submission and exists mainly to bring you back for another attempt. The tips box has an obvious incentive to optimise for the next score on this tool specifically, rather than for anything that would generalise, and that incentive does not disappear just because the tips are more fluently written than they used to be.
Cross-tool comparison features. A tool that offers to show you how your result compares to what another named tool would have given is addressing a real reader need - working out whether a gap between two tools is meaningful. But a genuine conversion between two tools' scales would require both to be stable and both populations to be comparable, conditions that rarely hold, so a feature promising this is either doing something more modest than it sounds like, or making a claim it cannot support.
What decides which bucket a given feature lands in
The test is not whether a feature is new or whether a tool calls it an improvement. It is whether the feature gives you something to check - a component to attribute a change to, a stored condition to compare against, a documented population - or whether it only changes how the same underlying number is presented. What the underlying models are capable of next will keep expanding what a tool could show; whether a given tool chooses to show the comparability-improving version of a feature or the engagement-improving version remains a product decision each time, unrelated to the model's capability.
None of these trends touch the two properties that stay outside this category by design: a physical measurement does not get more axes or more prose, and a human reviewer's response was never structured around any of this to begin with. Rate Cock already sits on the comparability side of several of these - six independently reported axes and visible breakdowns on public entries - which is a reasonable baseline for judging whether a newer feature elsewhere in the category is adding information or just adding motion.