Scores
The table people want, and why nobody can publish it
A chart mapping tool A's 7 to tool B's 8 would need both scales to be stable and both populations to be the same; neither holds.
Guides on Scores: How much weight a rating tool result deserves, Every part of a rating result, and what each one is for, What can and cannot be compared, and how
No reliable conversion table between rating tools exists, and none can. A table needs both tools' scales to stay fixed and their user populations to be comparable, and neither holds for long - the problem is structural, not a gap nobody has filled.
What the table would require
A stable conversion needs two things to hold still: each tool's scale has to be fixed, and the two tools' populations have to be comparable enough that a percentile on one means roughly the same thing as a percentile on the other.
Neither holds in practice. Rubrics get edited and mappings get retuned without a public changelog, so a conversion accurate today can be wrong after either tool's next quiet update. Quiet updates are measurable even in flagship models: Chen, Zaharia and Zou (2023) found GPT-4's accuracy on one task fell from 84% to 51% between its March and June 2023 versions. And the two tools are drawing from different pools of submitters, with different habits about which photo they choose to submit, which means "equivalent percentile" is itself a moving, unverifiable target.
Why the rank-first method survives where a table cannot
A published table tries to solve the comparison once and reuse it forever. Converting each score to its own tool's rank at the time of comparison solves it fresh each time, which is more work but does not go stale the way a static table does the moment either tool changes anything.
This is also why a tool's own architecture does not rescue the table idea: even two tools built on similar underlying models can diverge in their rubric and mapping enough that a fixed conversion between their outputs would still drift.
What this looks like elsewhere
A conversion table works fine for genuinely fixed units - inches to centimetres never drifts, because neither scale is a rubric that gets retuned. A measured figure converts cleanly for exactly that reason. A rating score is not that kind of quantity, and a human reviewer's opinion is further still - there is no scale behind a judge's read of a submission to build a table against in the first place. Rate Cock states plainly what a result on its own platform means; what your score means is written for one tool's own scale, and that is the honest scope any such explanation can have - one tool, not a bridge to every other one.
The table people want would need both scales to hold still long enough to draw it. Neither does, so the table stays undrawable, and the rank-first method is the substitute that actually works.