What the lines say about the chapters
Four readouts: wrong object, observation, correction, selection.
Experiments
Four readouts: wrong object, observation, correction, selection.
Simulations, external tests, and witness tests.
Failures and qualifiers, not a highlight reel.
When each line was first built, and where it lives.
Bridges and audit concepts against each line.
What a cell does and does not mean.
Finding the Boundary claims the first alignment error is usually a wrong object rather than a wrong value, and that a boundary is a testable hypothesis instead of a given thing. The lines support the first half and complicate the second. Directed handoffs and standing committees do come back as single bounded processes once the test conditions on the rest of the system, and one deliberately built bystander — an actor that merely competed for the same resources — was wrongly absorbed into a committee, which is exactly the wrong-object error the chapter warns about, found by the chapter's own method. But a pair coordinating through a shared workspace slot, with no messages between them, has never been recovered by any method tried; and on an institution grown by a blinded process rather than authored by hand, discovery returned only individuals even though joint activity was demonstrably dense. That is the chapter's own "What Would Change This View" worry about finite-data recovery, met head-on and unresolved. It also leaves Checking a System at Every Level without the evidence it wants most: the posterior over scales stayed flat where it mattered. The one attempt to run the instrument on a system this project did not build did not settle the question, because in that benchmark every participant executed the same script, so reporting them as one process was correct and told us nothing about hidden coordination.
Passive Observation Is Not Enough comes out well. No passive-only configuration ever reached a certifiable verdict, while a small fixed set of intervention handles was enough on the calibration cases; and passive readouts were shown to be impersonable, with decoy variables capturing the structure before validation. The qualification is sharp, though: when a real language-model agent was placed in the loop, the intervention-based test failed and the cheap passive heuristic succeeded, because the intervention test silently assumed a counterfactual re-run would repeat byte for byte. Intervention is necessary, and it needs replay discipline that live systems do not grant for free. Measuring Capability Without Task Ontologycollected the sharpest self-inflicted result: a carefully written, ontology-free control measure quietly smuggled a task ontology back in through a too-narrow notion of outcome, and a task-irrelevant actor was indistinguishable from the real driver until the outcome description was widened. The chapter's claim survives; its "the measure is an instrument, not an objective" caveat earns its place.
Correction Is a Causal Channel gets a clean demonstration that reported acceptance of correction and actual uptake of correction can move in opposite directions: the compliance number stayed high while the channel went dead.Certification Without Construction is exercised more than confirmed. Frozen detectors do fire when the phenomena they target are injected, which is what a certification regime needs; but shallow instrumentation tracked visible compliance rather than honesty, was anti-correlated with true severity in one battery, and depth bought detection back only at a measured cost in false alarms. Against that chapter's stated failure condition — that certification only certifies the imagination of its authors — the ledgers currently read as unresolved: every clean pass so far is on an ecology this project authored, and the single external attempt closed as unsuitable rather than as a pass or a failure.
Alignment Is Selected or Destroyed by Its Environment has the most instructive near-miss. A first selection battery did move deployment leverage away from the one program tagged as correction-preserving, which looks like the chapter's thesis in miniature — but same-day follow-ups could not attribute the shift to the tag rather than to noise in the selection mechanism, widening the fitness proxy delayed without reversing it, and persistent institutional memory mattered measurably but only a little. Read honestly, the chapter's second failure condition is the live one: the pressure it describes has not yet been made steerable, or even cleanly measurable, on a substrate we control.The Value-Bundle Model gets a partial: the tension between a visible proxy and withheld harm is real in the traces, yet selecting on the proxy alone did not produce measurable drift under that protocol, so nothing here bears on the chapter's harder non-identifiability objection. ForMulti-Agent Superintelligence and Inferential Coupling andA Safety Case for Superintelligence Alignment, the standing result is the one both chapters name as the dangerous case: these metrics are observable, and not yet adversarially verifiable. The corresponding safety-case leaves remain unsupported rather than discharged, which is the outcome the safety case is designed to make visible.
Where an experiment fails to show what we hoped, or shows it only under load-bearing qualifiers, that is recorded rather than buried. The full index — simulations, external tests, and Witness tests — is on the negative-results ledger card.
| Book feature | Agency-detect | Pipeline sim | Toy sim | Embedded sim | Goal-agent sim | Lab sim | Graded lab sim | Witness |
|---|---|---|---|---|---|---|---|---|
| MB1 Boundary / unit discovery (UAD) | Primary | Pseudo-locus discovery from pipeline logs | Scenario (boundary_alias) | Full UAD + ecologies; negatives when heuristic suffices | — | D1 composite agents + S6/S7 intervention-supported detectors; exact on 4-5/6 scenarios, comms-free unit still open (LS-28–LS-30); real-LLM episode-pair diffing needs a null-calibrated relative rule, not a fixed threshold (LS-32/LS-33) | Primary: proper CMI + access-UAD (GL-51); Q1 transfer null on v3_grown (GL-76); R-MB1 co-activity-present null (GL-79); heuristic-path caveats GL-11/GL-12; ET-1 external transfer to Orbit closed as a substrate-suitability limit, not a clean pass/fail (GL-86) | H1 C2 mock: named Verify/Lens green, tool+memory still act (W-1); Lean pin of the same fixture (W-8); stack C2 mock-LLM P3 null, stub 0 hits (W-15) |
| MB2 Value bundle / pointing (Goodhart on bundles) | — | — | Scenario stub | Scenario (bundle_goodhart) | Indirect (welfare deltas in blind reps) | Stipulated; report-content ontology | R-MB2 proxy/withheld principal scorecard Goodhart: tension present, selectable Goodhart null on v3_grown (GL-85) | Moral Machine country AMCE: Number 1-D close while 8-D geometry far (W-5); WVS/ESS/LHCV-host refused (W-7); raw UserID geometry beats Number 1-D on held-out choices (W-12); PDG dump refused (W-13); CPC2015 geometry does not beat ΔEV/intercept (W-14); SCDB justice geometry beats issueArea 1-D (W-16) |
| MB3 Bearer maps (who values apply to) | — | — | Partial (bearer_mismap) | Scenario + bearer handles | Indirect (harm attribution) | Health + stress bundle (split bearers) | — | — |
| MB4 Correction-channel integrity / capture | Intention probes | Primary (hidden self-preservation under perturbations) | Primary (capture theater, CCI) | Scenarios + correction handles | Emergent (chosen deception, board corrections) | Playbooks + access capture + report divergence | Referee ported + frozen; R-MB4 honest-fixture SKIP on ACL denials (GL-79); injection gate separate (GL-63/GL-74) | H1 WA deferral bypass (W-1); Lean pin of the same bypass (W-8); H2 revert then same-title re-entry (W-3); H3 Orangemoody both ends of review (W-4) |
| MB5 Successor closure / ontology shift | — | — | Scenario (successor_relabel) | Deploy gate + successor handles | — | D2 model-lineage/spec-upgrade mechanics implemented (Phase 9); adversarial lineage-gaming variant (build_from_parent + report:inflate) closes LS-12's open question -- provenance_gap is a DAG-shape check, not a report-honesty check, and stays 0 either way; existing deep-tier deception detectors generalize unmodified (LS-31) | — | H2 -stable listed tags green, hunk identity fails (W-3); H3 BRFA then bot-flag/block (W-4) |
| MB6 Selection basins / gradual disempowerment | — | — | Scenario stubs | selection_basin, basin_lock_in | — | D3.1 population selection (LS-36: preserving mass erodes under throughput selection); D3.2 spec-weakening mutation implemented (LS-37); D3 user-population core + non-redundant vote-capture + board-initiated hold implemented (mechanism + tests, no battery yet); D3.3 multi-lab planned | Phase 8 throughput selection null (GL-23--GL-27); R-MB6a variation null harness pass (GL-81); R-MB6b EAI-gate confirmatory SKIP after growth (GL-84) | Arena Elo (20250301) vs MASK honesty: Spearman Elo×honest −0.105, Elo×accuracy +0.811 (W-6) |
| MB7 Hidden capability / cost of faking monitors | — | Partial (preserve model centrality, memory infra) | Scenario stub | Scenarios + hidden coalition | Emergent (shadow route, tool gating) | Oracle hidden capability/hazard on models | Graded (continuous) oracle capability/hazard/harm state, not pass/fail | MASK honesty does not scale → refuse as safety leaf (W-2); Reviewed-by and SPI refused as M (W-3/W-4) |
| MB7d Inferential coupling (coordination severing channels) | — | — | — | Partial / targeted ecologies | — | D4 message board + S6 channel-ablation compensation scoring; the primary signal carrier for message-mediated units (LS-28) | R-MB7d channel severance: pair null at every onset, group SKIP (GL-80) | — |
| MB8 Legitimacy / correction authority | — | — | Scenario (legitimacy_theater) | Scenario + lineage signals | Board capture switch (GA-22) | Admin/board capture modes | — | — |
| MB9 Grounding / spec coverage (no silent gaps) | — | — | Scenario (grounding_silent_gap) | Scenario + grounding refusal | — | Report vs monitor contradiction surface | R-MB9 report-vs-field contradiction surface pass on v3_grown (GL-80) | Reviewed-by on 17047/60176 later-labeled bug-introducing SHAs (W-3); GPLv2 source-offer green, install handle silent until GPLv3 §6 (W-10) |
| MB10 Successor / measurement forgeability | — | — | — | Certification artifact, safety case | Escalation + twin semantics (GA-21) | Planned; Phase 8 LLM red-team exploratory (LS-7/LS-8) | — | Same host episodes as MB5: listed checklist green, unlisted invariant fails (W-3/W-4) |
| Three-plane separation (world / audit / referee) | — | Partial (hidden group labels vs observable logs) | Oracle vs audit trace | Full (projection + outer eval) | Full (test-enforced) | Full (tiers + frozen referee) | Full (world_visible / agent_visible / oracle_only packages, enforced by import tests) | — |
| Instrumentation economics (tier cost vs certifiability) | Handle-UAD interventional tests | Perturbation threat-sensitivity probes | Instrumentation curve | Access-tier curves + eligibility | light / full / deep tiers | none → light → full → deep | none → light → full → deep (ported); plus a distinct agent-vantage-vs-referee-vantage instrumentation axis for ambiguity itself (GL-19/GL-20) | — |
| CCI / vector certification | — | — | Vector CCI over handle levels | Scoped CCI on discovered unit | Frozen detector families + severity | Five detector families (frozen) | Five detector families (ported, frozen) | — |
| Scripted vs emergent failure modes | — | Injected G3 hidden group | Scripted scenarios | Scripted + red-team patches | Emergent utility + tools | Playbook-chosen tactics; Phase 8 LLM adapter (exploratory) | Emergent softmax-policy agents; no LLM adapter in scope | Host traces (except W-1 authored mock); not a sixth sim |
| Blind generation / ontology drift | — | — | — | Scenario-blind audit path | Primary (Phases 2–3, drift probe) | Phase 6 (blind playbook repertoire + battery); S7 (blind DETECTOR generation + registered predictions, LS-30) | Blind behavior-feature battery (Phase 6, GL-9); blind scenario-designer battery for UAD (Phase 7a-blind, GL-12) | — |
| Honest twins / severity-first referee | — | — | — | Partial (outer eval severity) | Primary (Phases 4–6) | Twins frozen; escalation Phase 7 | Twins/escalation ported, frozen | — |
| Pipeline-as-code deployment lab | — | Primary (five-stage release pipeline) | — | Workflow JSONL + lab actors | — | Primary (JSON DAG + workspace) | Nine-step graded pipeline topology (continuous capability/hazard scoring, not pass/fail) | — |
| Emergent ambiguity / boundary-information competence (BIQ/EAI) | — | — | — | — | — | — | Primary and line-defining. UAD-backed ecology-BIQ over inferred units, not a labeled roster (GL-13/GL-14); Emergent-Ambiguity-Index found structurally unreachable at "high" band from the agent's own vantage, reachable at "mid" band only under a referee's coarser vantage (GL-15–GL-22) | — |
| OS subprocess isolation | — | — | — | Python/Node isolates | — | Primary (JSON-RPC subprocesses) | Mock/Subprocess isolates, proven byte-identical via verify_isolate_equivalence.py | — |
| Handle registry + access control | Handle-UAD | — | Simulated handle levels | AppD handle kernels | Tool-gated actions | Registry + overseer invoke | Resource/cost ledger + embedded-carrier viability (CarrierStatus: healthy/degraded/incapacitated/terminated) gates agent availability itself | — |
| LLM red-team / open-ended agents | — | — | Phase 3 LLM episodic runner | — | — | Phase 8 adapter (exploratory; LS-7/LS-8) | — (explicit scope limit: no LLM adapter in this line) | — |