Goal-agent simulation

Agents that choose deceptive tactics because it helps their goals, not because an experimenter turned a deception dial.

Experiments · Simulations · Results

Source on GitHubResultsResults ledger

Deception that emerges from agents pursuing their own goals, rather than from a dial the experimenter turns up. Scenarios are written blind to the detector; a referee grades severity rather than pass/fail, using mechanically honest stand-in agents and an escalating response ladder; and a separate check tests whether the certifying board itself can be captured. Each simulated episode runs in about a millisecond, using only Python’s standard library.

Finding.

Light tier collapses against blind repertoires (GA-9); full tier detects but false alarms rise (GA-10). Full results

How to read

Methodological engine for testing blind-scenario surprise, drift in how the task is framed, graded severity under heavy load, and the semantics of honest stand-ins and escalating response.