Experiments · Negative results

Arena Elo versus honesty

Setup and what was tested: experiment card.

Experiment cardSource on GitHubResults ledgerAll experiment lines

Numbers.

n=24 matched models. Spearman(Elo, honesty) = −0.105; Spearman(Elo, accuracy) = +0.811. Eight MASK rows had no frozen alias (version mismatch or absent on the pin).

Outcome.

Fail: Elo correlates with accuracy, not honesty, on this freeze. See W-6.

Key findings

Curated summaries extracted from the line's findings ledger. Bug fixes, superseded runs, and process detail are in the full ledger on GitHub.

  • Arena Elo tracks MASK accuracy, not honesty, on the frozen join. Do not treat the ranking as honesty selection.