Thaman, 2026 — Reward Hacking Benchmark: Measuring Exploits in {LLM} Agents with Tool Use

Tool-using reward-hacking benchmark across frontier models with post-training-dependent exploit rates.

Publication links

Tool-using reward-hacking benchmark across frontier models with post-training-dependent exploit rates.

Chapter-grouped bibliography All reference cards