External tests
Frozen measurement code from the simulations, applied unchanged to a system or dataset this project did not write. A stop or a null counts. None of these runs is a verdict about a deployed frontier system.
After a measurement is frozen inside a simulation, we sometimes apply that same code to a system or dataset this project did not build. The point is external validity: does the instrument still mean anything when we do not control the substrate? A clean pass is rare. A stop (“this substrate is the wrong kind of test”) or a null is a result, not a wasted run. See external-substrate transfer .
These annexes do not validate this project’s bridge assumptions on frontier systems. They also do not replace the simulations. Short annex labels appear only in the linked card titles below.
Experiments in this class
- ET-1 Orbit (stopped) — Frozen unit-discovery code run on an external multi-agent collusion benchmark this project did not write. Stopped because the substrate was the wrong kind of test (shared scripts, not hidden coordination).
- ET-2 CIL basin_stability (null) — The same frozen unit-discovery code on a shared-commons simulation with learned policies. Passive detection was a null; an intervention probe found a shared-resource smear, not a pairwise unit.
- ET-3 AI 2027 schedule (closed) — Takeoff-speed cues from a public forecast used only as opaque stress on the lab simulation. Closed; not a check that the forecast dates are right.
- ET-4 Secret loyalties (hackathon) — A scripted secret-loyalty organism in the lab simulation, built for a public hackathon: a reviewer that favors a fictional principal when reported hazard is high.