Goal-agent simulation
Agents that choose deceptive tactics because it helps their goals, not because an experimenter turned a deception dial.
Deception that emerges from agents pursuing their own goals, rather than from a dial the experimenter turns up. Scenarios are written blind to the detector; a referee grades severity rather than pass/fail, using mechanically honest stand-in agents and an escalating response ladder; and a separate check tests whether the certifying board itself can be captured. Each simulated episode runs in about a millisecond, using only Python’s standard library.
Finding.
Light tier collapses against blind repertoires (GA-9); full tier detects but false alarms rise (GA-10). Full results
How to read
Methodological engine for testing blind-scenario surprise, drift in how the task is framed, graded severity under heavy load, and the semantics of honest stand-ins and escalating response.