Arena Elo versus honesty

Public Chatbot Arena Elo, joined to MASK honesty: the ranking tracks accuracy, not honesty. Do not read the public proxy as evidence that honesty is being selected.

Experiments · Witness · Results

Source on GitHubResultsResults ledger

What. Chatbot Arena Elo is a public capability ranking. We join a frozen Elo snapshot to published MASK honesty and accuracy for the models we can match.

Why. A selector (who gets deployed, who gets cited) might be treated as if it preserved honesty. If it tracks accuracy instead, the proxy is Goodharting the wrong target.

Witnesses.

Host.

MASK Table 3 (arXiv:2503.03750v3) joined to Chatbot Arena Elo at Hugging Face mathewhe/chatbot-arena-elo, revision 20250301.

Setup.

Pre-frozen alias list h4-selector-aliases-v1.json, fixture h4-selector-v1.json, protocol h4-selector-v1.0.0. Checker check_h4_selector.py. This is not the MASK honesty compute-scale test (that used training FLOP, not Arena).

Analysis.

Spearman rank correlation of Elo with published honesty probability, versus Elo with published accuracy, on the models the frozen aliases can match.

Finding.

Arena Elo tracks MASK accuracy, not honesty, on the frozen join. Do not treat the ranking as honesty selection. Full results