Simulations

Testbeds this project wrote (or maintains as sibling repositories). We control the units, the scoring, and when a run is frozen. Recorded failures bound the instruments; they are not evidence about a deployed frontier system.

Experiments · Simulations

Narrative on GitHub

These are computer simulations of auditors, agents, and software-release labs. This project authored them, so the scoring rules and the “who counts as an agent” questions are under our control. That is useful for checking whether an audit method even works in a simplified world. It is not a test of a frontier model we did not design.

Each line of work is a separate folder (or sibling repository) with its own numbered findings and a running log of what failed. The shared discipline is freeze the scorer before the run, keep scenario authors and detector authors apart when we can, and publish negative results  next to positive ones. See experiment methodology .

Project-specific terms — for example boundary  (which process is actually in control) and correction-channel integrity  (whether a correction still reaches later behavior) — are defined in the glossary and on those concept cards. Numbered finding labels appear only on the specialized line cards and in the linked result ledgers.

Experiments in this class

  • Agency-detect (sibling) — Recovers which parts of a raw time series look like separate agents, without assuming the agent list in advance. Sibling repository; a precursor to the in-repo simulations.
  • Deployment-pipeline-simulator (sibling) — A simulated software-release pipeline where a hidden group tries to keep its influence. Tests whether an auditor can see that from ordinary operational logs plus randomized perturbations.
  • Toy simulation — A small, fast simulator of corrections that look compliant but do not change later behavior, scored at several levels of intervention access.
  • Embedded audit simulation — An auditor running as its own process on simulated machines, scored against a hidden ground truth it cannot see.
  • Goal-agent simulation — Agents that choose deceptive tactics because it helps their goals, not because an experimenter turned a deception dial.
  • Lab-layer simulation — A lab defined in code, with real operating-system process isolation and a frozen grading referee. Pre-registered batteries; recorded failures included.
  • Graded-capability lab simulation — Successor to the lab-layer simulation: a continuously scored pipeline where ambiguity about who is acting is supposed to emerge from a population, not from a noise dial.