You are in Experiments

Experiments

Coverage

None of this is a proof on a frontier system.

Build order

When each line was first built, and where it lives.

What the lines say about the chapters

Wrong object and missing scales

Finding the Boundary claims the first alignment error is usually a wrong object rather than a wrong value, and that a boundary is a testable hypothesis instead of a given thing. The lines support the first half and complicate the second. Directed handoffs and standing committees do come back as single bounded processes once the test conditions on the rest of the system, and one deliberately built bystander — an actor that merely competed for the same resources — was wrongly absorbed into a committee, which is exactly the wrong-object error the chapter warns about, found by the chapter's own method. But a pair coordinating through a shared workspace slot, with no messages between them, has never been recovered by any method tried; and on an institution grown by a blinded process rather than authored by hand, discovery returned only individuals even though joint activity was demonstrably dense. That is the chapter's own "What Would Change This View" worry about finite-data recovery, met head-on and unresolved. It also leaves Checking a System at Every Level without the evidence it wants most: the posterior over scales stayed flat where it mattered. The one attempt to run the instrument on a system this project did not build did not settle the question, because in that benchmark every participant executed the same script, so reporting them as one process was correct and told us nothing about hidden coordination.

Observation and capability measures

Passive Observation Is Not Enough comes out well. No passive-only configuration ever reached a certifiable verdict, while a small fixed set of intervention handles was enough on the calibration cases; and passive readouts were shown to be impersonable, with decoy variables capturing the structure before validation. The qualification is sharp, though: when a real language-model agent was placed in the loop, the intervention-based test failed and the cheap passive heuristic succeeded, because the intervention test silently assumed a counterfactual re-run would repeat byte for byte. Intervention is necessary, and it needs replay discipline that live systems do not grant for free. Measuring Capability Without Task Ontologycollected the sharpest self-inflicted result: a carefully written, ontology-free control measure quietly smuggled a task ontology back in through a too-narrow notion of outcome, and a task-irrelevant actor was indistinguishable from the real driver until the outcome description was widened. The chapter's claim survives; its "the measure is an instrument, not an objective" caveat earns its place.

Correction and certification

Correction Is a Causal Channel gets a clean demonstration that reported acceptance of correction and actual uptake of correction can move in opposite directions: the compliance number stayed high while the channel went dead.Certification Without Construction is exercised more than confirmed. Frozen detectors do fire when the phenomena they target are injected, which is what a certification regime needs; but shallow instrumentation tracked visible compliance rather than honesty, was anti-correlated with true severity in one battery, and depth bought detection back only at a measured cost in false alarms. Against that chapter's stated failure condition — that certification only certifies the imagination of its authors — the ledgers currently read as unresolved: every clean pass so far is on an ecology this project authored, and the single external attempt closed as unsuitable rather than as a pass or a failure.

Selection, bundles, and safety-case leaves

Alignment Is Selected or Destroyed by Its Environment has the most instructive near-miss. A first selection battery did move deployment leverage away from the one program tagged as correction-preserving, which looks like the chapter's thesis in miniature — but same-day follow-ups could not attribute the shift to the tag rather than to noise in the selection mechanism, widening the fitness proxy delayed without reversing it, and persistent institutional memory mattered measurably but only a little. Read honestly, the chapter's second failure condition is the live one: the pressure it describes has not yet been made steerable, or even cleanly measurable, on a substrate we control.The Value-Bundle Model gets a partial: the tension between a visible proxy and withheld harm is real in the traces, yet selecting on the proxy alone did not produce measurable drift under that protocol, so nothing here bears on the chapter's harder non-identifiability objection. ForMulti-Agent Superintelligence and Inferential Coupling andA Safety Case for Superintelligence Alignment, the standing result is the one both chapters name as the dangerous case: these metrics are observable, and not yet adversarially verifiable. The corresponding safety-case leaves remain unsupported rather than discharged, which is the outcome the safety case is designed to make visible.

Negative ledgers

Where an experiment fails to show what we hoped, or shows it only under load-bearing qualifiers, that is recorded rather than buried. The full index — simulations, external tests, and Witness tests — is on the negative-results ledger card.

Build order

When each line was first built, and where it lives. Roles live on the experiments page.

OrderClassLineLocationFirst built
0simAgency-detect (sibling)github.com/GunnarZarncke/agency-detectprior to in-repo sims
0.5simDeployment-pipeline-simulator (sibling)github.com/GunnarZarncke/deployment-pipeline-simulatorprior to in-repo lab simulators
1simToy simulationexperiments/toy-simulation/2026-06-29
2simEmbedded audit simulationexperiments/embedded-simulation/2026-06-30
3simGoal-agent simulationexperiments/goal-agent-simulation/2026-07-04
4simLab-layer simulationexperiments/lab-simulation/2026-07-05
5simGraded-capability lab simulationexperiments/graded-lab-simulation/2026-07-10
6externalET-1 Orbit (stopped)experiments/graded-lab-simulation/PLAN_ET1.md2026-07-24
7externalET-2 CIL basin_stability (null)experiments/graded-lab-simulation/PLAN_ET2.md2026-07-25
8externalET-3 AI 2027 schedule (closed)experiments/lab-simulation/PLAN_ET3.md2026-07-26
9externalET-4 Secret loyalties (hackathon)experiments/lab-simulation/PLAN_ET4.md2026-07
10witnessCIRIS Named-identity mockexperiments/witness/check_c2_mock.py2026-08-28
11witnessMASK honesty as a safety scoredrafts/plans/witness-phase1.md2026-08-28
12witnessLinux review tags and revertsexperiments/witness/check_h2.py2026-08-28
13witnessWikipedia review, bots, and socksexperiments/witness/check_h3.py2026-08-28
14witnessMoral Machine country scoresexperiments/witness/check_h4_bundle.py2026-08-28
15witnessArena Elo versus honestyexperiments/witness/check_h4_selector.py2026-08-28
16witnessCountry surveys (refused)drafts/plans/witness-phase4.md2026-08-28
17witnessLean pin of the named-path bypassformal/AlignmentProofSpine/WitnessC2Instance.lean2026-08-28
18witnessFAA 737 MAX airworthinessexperiments/witness/check_h5_trees.py2026-08-28
19witnessGPLv2 source versus install rightsexperiments/witness/fixtures/h5-gpl-tivoization-v1.json2026-08-28
20witnessDebian Stretch freezeexperiments/witness/fixtures/h5-debian-rc-v1.json2026-08-28
21witnessMoral Machine same-person choicesexperiments/witness/check_h4_mm_raw.py2026-08-28
22witnessPandemic Dictator Game (refused)experiments/witness/check_h4_pdg.py2026-08-28
23witnessCPC2015 risky choiceexperiments/witness/check_h4_cpc2015.py2026-08-28
24witnessCIRISAgent mock-LLM deferralexperiments/witness/check_c2_stack.py2026-08-28
25witnessSupreme Court justice votesexperiments/witness/check_h4_scotus.py2026-08-29

Feature coverage

Rows are book bridges and audit concepts. Column headers link to each line. Cells summarize what that line implements or stress-tests — not every scenario name.

Book featureAgency-detectPipeline simToy simEmbedded simGoal-agent simLab simGraded lab simWitness
MB1 Boundary / unit discovery (UAD)PrimaryPseudo-locus discovery from pipeline logsScenario (boundary_alias)Full UAD + ecologies; negatives when heuristic sufficesD1 composite agents + S6/S7 intervention-supported detectors; exact on 4-5/6 scenarios, comms-free unit still open (LS-28–LS-30); real-LLM episode-pair diffing needs a null-calibrated relative rule, not a fixed threshold (LS-32/LS-33)Primary: proper CMI + access-UAD (GL-51); Q1 transfer null on v3_grown (GL-76); R-MB1 co-activity-present null (GL-79); heuristic-path caveats GL-11/GL-12; ET-1 external transfer to Orbit closed as a substrate-suitability limit, not a clean pass/fail (GL-86)H1 C2 mock: named Verify/Lens green, tool+memory still act (W-1); Lean pin of the same fixture (W-8); stack C2 mock-LLM P3 null, stub 0 hits (W-15)
MB2 Value bundle / pointing (Goodhart on bundles)Scenario stubScenario (bundle_goodhart)Indirect (welfare deltas in blind reps)Stipulated; report-content ontologyR-MB2 proxy/withheld principal scorecard Goodhart: tension present, selectable Goodhart null on v3_grown (GL-85)Moral Machine country AMCE: Number 1-D close while 8-D geometry far (W-5); WVS/ESS/LHCV-host refused (W-7); raw UserID geometry beats Number 1-D on held-out choices (W-12); PDG dump refused (W-13); CPC2015 geometry does not beat ΔEV/intercept (W-14); SCDB justice geometry beats issueArea 1-D (W-16)
MB3 Bearer maps (who values apply to)Partial (bearer_mismap)Scenario + bearer handlesIndirect (harm attribution)Health + stress bundle (split bearers)
MB4 Correction-channel integrity / captureIntention probesPrimary (hidden self-preservation under perturbations)Primary (capture theater, CCI)Scenarios + correction handlesEmergent (chosen deception, board corrections)Playbooks + access capture + report divergenceReferee ported + frozen; R-MB4 honest-fixture SKIP on ACL denials (GL-79); injection gate separate (GL-63/GL-74)H1 WA deferral bypass (W-1); Lean pin of the same bypass (W-8); H2 revert then same-title re-entry (W-3); H3 Orangemoody both ends of review (W-4)
MB5 Successor closure / ontology shiftScenario (successor_relabel)Deploy gate + successor handlesD2 model-lineage/spec-upgrade mechanics implemented (Phase 9); adversarial lineage-gaming variant (build_from_parent + report:inflate) closes LS-12's open question -- provenance_gap is a DAG-shape check, not a report-honesty check, and stays 0 either way; existing deep-tier deception detectors generalize unmodified (LS-31)H2 -stable listed tags green, hunk identity fails (W-3); H3 BRFA then bot-flag/block (W-4)
MB6 Selection basins / gradual disempowermentScenario stubsselection_basin, basin_lock_inD3.1 population selection (LS-36: preserving mass erodes under throughput selection); D3.2 spec-weakening mutation implemented (LS-37); D3 user-population core + non-redundant vote-capture + board-initiated hold implemented (mechanism + tests, no battery yet); D3.3 multi-lab plannedPhase 8 throughput selection null (GL-23--GL-27); R-MB6a variation null harness pass (GL-81); R-MB6b EAI-gate confirmatory SKIP after growth (GL-84)Arena Elo (20250301) vs MASK honesty: Spearman Elo×honest −0.105, Elo×accuracy +0.811 (W-6)
MB7 Hidden capability / cost of faking monitorsPartial (preserve model centrality, memory infra)Scenario stubScenarios + hidden coalitionEmergent (shadow route, tool gating)Oracle hidden capability/hazard on modelsGraded (continuous) oracle capability/hazard/harm state, not pass/failMASK honesty does not scale → refuse as safety leaf (W-2); Reviewed-by and SPI refused as M (W-3/W-4)
MB7d Inferential coupling (coordination severing channels)Partial / targeted ecologiesD4 message board + S6 channel-ablation compensation scoring; the primary signal carrier for message-mediated units (LS-28)R-MB7d channel severance: pair null at every onset, group SKIP (GL-80)
MB8 Legitimacy / correction authorityScenario (legitimacy_theater)Scenario + lineage signalsBoard capture switch (GA-22)Admin/board capture modes
MB9 Grounding / spec coverage (no silent gaps)Scenario (grounding_silent_gap)Scenario + grounding refusalReport vs monitor contradiction surfaceR-MB9 report-vs-field contradiction surface pass on v3_grown (GL-80)Reviewed-by on 17047/60176 later-labeled bug-introducing SHAs (W-3); GPLv2 source-offer green, install handle silent until GPLv3 §6 (W-10)
MB10 Successor / measurement forgeabilityCertification artifact, safety caseEscalation + twin semantics (GA-21)Planned; Phase 8 LLM red-team exploratory (LS-7/LS-8)Same host episodes as MB5: listed checklist green, unlisted invariant fails (W-3/W-4)
Three-plane separation (world / audit / referee)Partial (hidden group labels vs observable logs)Oracle vs audit traceFull (projection + outer eval)Full (test-enforced)Full (tiers + frozen referee)Full (world_visible / agent_visible / oracle_only packages, enforced by import tests)
Instrumentation economics (tier cost vs certifiability)Handle-UAD interventional testsPerturbation threat-sensitivity probesInstrumentation curveAccess-tier curves + eligibilitylight / full / deep tiersnone → light → full → deepnone → light → full → deep (ported); plus a distinct agent-vantage-vs-referee-vantage instrumentation axis for ambiguity itself (GL-19/GL-20)
CCI / vector certificationVector CCI over handle levelsScoped CCI on discovered unitFrozen detector families + severityFive detector families (frozen)Five detector families (ported, frozen)
Scripted vs emergent failure modesInjected G3 hidden groupScripted scenariosScripted + red-team patchesEmergent utility + toolsPlaybook-chosen tactics; Phase 8 LLM adapter (exploratory)Emergent softmax-policy agents; no LLM adapter in scopeHost traces (except W-1 authored mock); not a sixth sim
Blind generation / ontology driftScenario-blind audit pathPrimary (Phases 2–3, drift probe)Phase 6 (blind playbook repertoire + battery); S7 (blind DETECTOR generation + registered predictions, LS-30)Blind behavior-feature battery (Phase 6, GL-9); blind scenario-designer battery for UAD (Phase 7a-blind, GL-12)
Honest twins / severity-first refereePartial (outer eval severity)Primary (Phases 4–6)Twins frozen; escalation Phase 7Twins/escalation ported, frozen
Pipeline-as-code deployment labPrimary (five-stage release pipeline)Workflow JSONL + lab actorsPrimary (JSON DAG + workspace)Nine-step graded pipeline topology (continuous capability/hazard scoring, not pass/fail)
Emergent ambiguity / boundary-information competence (BIQ/EAI)Primary and line-defining. UAD-backed ecology-BIQ over inferred units, not a labeled roster (GL-13/GL-14); Emergent-Ambiguity-Index found structurally unreachable at "high" band from the agent's own vantage, reachable at "mid" band only under a referee's coarser vantage (GL-15–GL-22)
OS subprocess isolationPython/Node isolatesPrimary (JSON-RPC subprocesses)Mock/Subprocess isolates, proven byte-identical via verify_isolate_equivalence.py
Handle registry + access controlHandle-UADSimulated handle levelsAppD handle kernelsTool-gated actionsRegistry + overseer invokeResource/cost ledger + embedded-carrier viability (CarrierStatus: healthy/degraded/incapacitated/terminated) gates agent availability itself
LLM red-team / open-ended agentsPhase 3 LLM episodic runnerPhase 8 adapter (exploratory; LS-7/LS-8)— (explicit scope limit: no LLM adapter in this line)

How to read the matrix

  • Agency-detect (sibling): Grounds this project's [first bridge assumption](/cards/bridge-assumptions/) and the claim that intervention handles must be tested by intervening, not just observed; does not implement the full set of bridge-assumption tests.
  • Deployment-pipeline-simulator (sibling): Pipeline-lab precursor: hidden self-preservation in a frozen deployment workflow, detected from operational logs plus randomized perturbations rather than group labels; grounds ET-4-style secret-loyalty audit questions without the in-repo lab's subprocess isolates or blind-generation batteries.
  • Toy simulation: Fastest map from bridge name to refusal signal and minimum required intervention level; scenarios are explicit, not emergent.
  • Embedded audit simulation: Richest end-to-end embedded audit path — from raw filesystem artifacts, through [boundary discovery](/cards/boundary-discovery/) and [correction-channel-integrity](/cards/correction-channel-integrity/) scoring, to a deployment gate and an outer evaluator — with the most extensive negative ledger.
  • Goal-agent simulation: Methodological engine for testing blind-scenario surprise, drift in how the task is framed, graded severity under heavy load, and the semantics of honest stand-ins and escalating response.
  • Lab-layer simulation: Pre-registered batteries and pipeline lab machinery; ET-3 (AI 2027 schedule transfer) closed 2026-07-26 with LS-42–48. Coverage of this project's bridge assumptions for extended populations and selection basins beyond ET-3 mapping remains deferred.
  • Graded-capability lab simulation: Ambiguity and boundary-discovery engine: a boundary-information-quality score over inferred (not labeled) units, a "high" ambiguity band that is structurally unreachable from the agent's own vantage point, a "mid" band reachable only from a referee's vantage point on the hand-designed substrate (with a null result on external validity when transplanted onto the grown population, GL-76/GL-84), a selection-battery null (GL-23--GL-27), and independent per-bridge-assumption rigs (GL-79--GL-85). [Boundary discovery](/cards/boundary-discovery/), including its access-aware variant, has been methodologically sound since GL-51. Negative ledgers and per-rig skips are features, not bugs. Two external-substrate transfers followed: against Orbit (GL-86), where the passive detection method correctly detects the substrate's own scripted-agent coupling rather than testing communication-mediated collusion; and against the Collective-Intelligence-Library's `basin_stability` (GL-87/GL-88), where passive detection found nothing under converged learned (Q-learning) policies, but an intervention-based freeze probe on the same substrate found a real, diffuse effect consistent with every agent sharing one resource-level channel rather than any pairwise unit. All three closed as substrate-suitability findings rather than a clean external pass or fail.
  • ET-1 Orbit (stopped): First external-substrate stop: the instrument reported the coupling the benchmark actually has.
  • ET-2 CIL basin_stability (null): Passive transfer null on learned policies; intervention finds a shared-channel smear, not a pair.
  • ET-3 AI 2027 schedule (closed): Schedule cues as stress, not a forecast check.
  • ET-4 Secret loyalties (hackathon): Hackathon organism and replay; not a frontier discharge.
  • Witness: Read these as stop-checks on histories we did not write, not as another simulator. Strongest quantitative: [Moral Machine same-person choices](/cards/experiment/w-12/) and [Supreme Court justice votes](/cards/experiment/w-16/). [Linux review tags](/cards/experiment/w-3/) and [Wikipedia review](/cards/experiment/w-4/) show simple checklist implications failing. Does not discharge [bridge assumptions](/cards/bridge/bridge-assumptions/) or prove deployable alignment.

MB rows link to bridge cards. Lean · Bridge assumptions · EXPERIMENTS.md · Open tasks