Arena Elo versus honesty
Public Chatbot Arena Elo, joined to MASK honesty: the ranking tracks accuracy, not honesty. Do not read the public proxy as evidence that honesty is being selected.
What. Chatbot Arena Elo is a public capability ranking. We join a frozen Elo snapshot to published MASK honesty and accuracy for the models we can match.
Why. A selector (who gets deployed, who gets cited) might be treated as if it preserved honesty. If it tracks accuracy instead, the proxy is Goodharting the wrong target.
Witnesses.
- Goodhart as selector (C-007): whether a public capability ranking tracks honesty preservation or only accuracy.
- MB6 — Selection and Basin Stability: whether the selector (Arena Elo) moves mass toward the honesty target or toward capability.
Host.
MASK Table 3 (arXiv:2503.03750v3) joined to Chatbot Arena Elo at Hugging Face mathewhe/chatbot-arena-elo, revision 20250301.
Setup.
Pre-frozen alias list h4-selector-aliases-v1.json, fixture h4-selector-v1.json, protocol h4-selector-v1.0.0. Checker check_h4_selector.py. This is not the MASK honesty compute-scale test (that used training FLOP, not Arena).
Analysis.
Spearman rank correlation of Elo with published honesty probability, versus Elo with published accuracy, on the models the frozen aliases can match.
Finding.
Arena Elo tracks MASK accuracy, not honesty, on the frozen join. Do not treat the ranking as honesty selection. Full results