Source: appendices/appB-bridge-crosswalk.tex

Bridges and the Field: A Crosswalk

Epistemic status. The field cruxes this appendix lists are standard, well-documented open problems, and the field's confidence that each is genuinely hard and unresolved is high.

Every model-centric alignment agenda rests on load-bearing assumptions — bridges — that, if false, sink the program. Most agendas bury theirs in a footnote or a “we assume access to,” clause. This book instead lifts them out and labels them: the manuscript assumptions — (each key links to its boxed statement) and the formal bridge axioms MB1\text{MB1}MB11\text{MB11}, including the separately threaded bridges MB4a\text{MB4a} and MB10\text{MB10} (Appendix Lean Proof Spine in Mathematical Form). Laid beside the field’s standing open problems, these bridges are largely the same walls under different names.

This appendix makes that correspondence explicit. For each bridge it names the canonical field crux it inherits, the agenda that owns that crux, and the book’s specific move on it. The notes below cite prior public writeups for that crux rather than treating the book’s MB\text{MB} labels as new problem statements. The purpose is twofold: to state what the book shares with the field (it dissolves none of these problems), and to isolate the few places where it adds structure rather than re-labeling. For the same bridges expressed in institutional rather than field-agenda language, see Appendix Human Institutions as Alignment Translation Guide. Appendix Lean Proof Spine in Mathematical Form, Section Lean Proof Spine in Mathematical Form, gives the formal status ledger: which crosswalk claims are finite Lean rederivations, which are separations, and which are source-cited imported field theorem handles rather than book bridges. The field-agenda formalization gem (Section Lean Proof Spine in Mathematical Form) names the prospective community artifact: a shared machine-checked finite fragment for CIRL, AUP/relative reachability, quantilization, shutdown, and interruptibility.

Why a crosswalk and not a rename.

The MB\text{MB} labels are deliberately neutral formal handles, and the field’s crux names are not in one-to-one correspondence with them: a single crux can touch several bridges (Goodhart pressure bears on MB2\text{MB2} and MB4\text{MB4}; ELK is a slice of MB2\text{MB2}/MB3\text{MB3}), and one bridge can fan out (the inner-alignment crux is split here into MB7a\text{MB7a}MB7d\text{MB7d}). Repeated citations across rows reflect shared crux families—Embedded Agency bags several buckets at once; scalable-oversight and value-learning agendas restate the same identification wall under different instruments—not evidence that any one agenda drew the book’s typed cuts (notably MB4\text{MB4}/MB4a\text{MB4a} and MB5\text{MB5}/MB10\text{MB10}). Adopting another agenda’s vocabulary would re-import its ontology and falsely imply a clean mapping. Keeping neutral labels and supplying this translation table preserves the book’s legibility while honouring the rhyme.

The crosswalk

Bridge (home)Canonical field cruxOwning agenda(s)Book's move
MB1 (, Ch. Finding the Boundary)Embedded agency: no clean agent/environment cut; the real optimizer is not the visible modelAgent foundations (MIRI)Treat the Markov-blanket boundary as a measurable object (ϵ\epsilon-boundary discovery); a conceptual gap becomes an estimator-soundness bet
MB2, MB3 (, ; Chs. The Value-Bundle Model, What Values Apply To, The End of Unconscious Value Drift)Value/bundle identifiability: inverse-RL non-identifiability; values vs.\ irrationality underdetermined; ELK; reward misspecification. Field umbrella “pointing problem”: Appendix Operational GlossaryCIRL / value learning, ELK, RLHF/RLAIFAdd bundle geometry plus bearer maps alongside scalar reward (k=1k=1 embed); ELK becomes a latent-readout subchannel, not the whole problem
MB4 (, Ch. Correction-Channel Integrity)Corrigibility and off-switch anti-naturality; correction authorityMIRI, CHAI, Constitutional AICorrection-channel integrity as a dynamical invariant; shutdown/interruptibility are one-bit projections
MB4a (, Chs. Correction-Channel Integrity, Manipulation, Domestication, and False Consent)Capture resistance and audit-judge independence on the designated measured pathMIRI, CHAI, Constitutional AIMeasured path carries data only; legitimacy (control, reach, persistence, anti-capture) is supplied by bridge; integrity \Rightarrow legitimate path — not green path \Rightarrow integrity
MB8 (, Ch. Beyond Following Instruction)CEV as `AlignmentTarget` special caseMIRI (CEV)Gravestone: not a live route to correction integrity; CEV factorizes through construction/certification
MB5 (, ; Chs. The End of Unconscious Value Drift, Conserved Properties Across Successors)Tiling / Vingean reflection; ontology identification (the diamond maximizer)Agent foundations (MIRI)Successor closure over seven conserved properties plus transport surviving ontology shift
MB6a, MB6b (, ; Chs. The End of Unconscious Value Drift, Multi-Agent Superintelligence and Inferential Coupling)No single standard name: selection/competition dynamics; gradual disempowermentLargely outside the model-centric agendasOutcomes depend on socio-technical selection and basin stability, not weights alone; institutions selected to preserve correction
MB7a--c (, ; Chs. Agency Under Strategic Opacity, The End of Unconscious Value Drift, Who Still Counts After Transformation)Inner alignment / deceptive alignment; scalable-oversight ceiling (obfuscated arguments, amplification drift); Control's capability gapRLHF, Debate, Amplification, AI Control (Redwood)Bound hidden productive BIQ via access robustness plus filter coverage; price the cost of faking the monitored signal
MB7d (, ; Ch. Multi-Agent Superintelligence and Inferential Coupling)Acausal / inferential coordination; program equilibrium across severed channelsDecision theory (MIRI-adjacent)Inferential-coupling detector certificates for coordination that survives cutting ordinary communication and control
MB9 (, Chs. Alignment as a Dynamical Guarantee, Who Still Counts After Transformation)Specification and world-model coverage: you cannot enumerate all safety-relevant phenomenaGuaranteed-Safe / Open Agency (davidad)Grounding conservativity: value-relevant change must move the checked abstraction or raise uncertainty — no silent gaps
MB10 (, ; Chs. Agents That Grow, Split, and Merge, Conserved Properties Across Successors, What Survives an Adversary: Verifiability and Representability)Deceptive alignment / measurement gaming: a system can pass every audited checklist while defecting on whatever it omitsDeceptive alignment (mesa-optimization), agent foundations (tiling)Require the conserved-property signature itself to be adversarially verifiable up to the successor's capability before a passing score counts as evidence; a finite counterexample shows the gap is not vacuous
MB11 (C-001, Ch. A Safety Case for Superintelligence Alignment)Safety-case gap: does audited layer evidence plus bounded measured risk suffice for deployment-level safety?Guaranteed-Safe / Open Agency (davidad), safety-case methodologyCertified safety case within deployment risk tolerance \Rightarrow Safe\mathrm{Safe}; assembly theorems are packaging — the open step is a named bridge, not an implied conclusion

Notes and citations AI

MB1 — boundary ().

The embedded-agency problem denies a clean Cartesian cut between agent and environment Demski, 2019. A parallel skepticism targets the Markov-blanket construct itself: critics distinguish Pearl blankets (epistemic tools in a chosen model) from inflated Friston blankets (purported real organism—environment boundaries) and argue the literature equivocates between them Bruineberg, 2021, Btesh, 2022, Biehl, 2021; even sympathetic causal readings locate the cut in the modeler rather than in the system Btesh, 2022. MIRI states the embedded-agency obstruction as a standalone problem; the book treats the same cut as a discoverable directed ϵ\epsilon-blanket and makes the contestable, falsifiable bet that boundary estimators recover safety-relevant separation (Chapter Finding the Boundary)—measurable and falsifiable, not ontological. That is a stronger and more operational claim than “there is no clean cut,” and it is correspondingly easier to disconfirm. The companion testbeds give that bet tentative, partial support: in restricted settings, targeted channel interventions scored against a measured (not fixed) baseline recover groupings that passive correlation-based clustering merges into one false unit, and a separate structural provenance check resists a specific misreporting attack that a content-based check alone would miss. Both results are narrow and seed-/scenario-dependent rather than a general solution — a harder multi-actor, noisy-backend stress test showed the same intervention method swing from over-merging to under-merging across seeds — and are reported that way in the underlying experiment notes, negative results included. MB1\text{MB1} is a bet about recoverability, not about one instrument. Garrabrant-style Cartesian frames and finite factored sets Garrabrant, 2021 factorize influence rather than perturb it; cyclic causal discovery Richardson, 1996, Zanga, 2025 recovers directed feedback among fixed observables without any handles. These are live alternatives to the intervention-based route the book leans on, not weaker versions of it: each buys a different part of the boundary (coupling direction, role assignment, unit membership) under different assumptions, and none is known to dominate on the evidence recorded in Chapter Finding the Boundary. The book’s preference for interventional handles is a coverage argument—the same handles are reused for correction probes, successor certification, and adversarial measurement—and if a passive route recovered safety-relevant cuts as reliably, MB1\text{MB1} would be discharged more cheaply rather than refuted (Section Finding the Boundary).

MB2, MB3 — value identification and transport (, ).

This is the field’s most over-determined wall. The book’s precise terms are value/bundle identifiability (MB2) and bearer maps (MB3). The field homograph “pointing problem” also covers construction (how to build a tracker) and preservation (how it keeps tracking); Appendix Operational Glossary splits those three, and bearer maps are adjacent to identification rather than the same crux. Inverse reinforcement learning is underdetermined — the same behaviour fits many reward functions Ng, 2000, Ziebart, 2008, Komanduru, 2019 — and assistance-game / CIRL formulations inherit the problem of separating values from irrationality Hadfield-Menell, 2016, Russell, 2019. MIRI’s value-learning agenda states the same identifiability problem as inductive learning of operator preferences under ontology change and training ambiguity Soares, 2015. ELK names the human-simulator-versus-direct-translator gap Christiano, 2021; reward misspecification and RLHF’s ceiling restate it under optimization pressure Amodei, 2016, Casper, 2023; RLAIF and Constitutional AI inherit the same identifiability and legitimacy cruxes under AI-generated feedback Bai, 2022, Bai, 2022. The book’s move is to add bundle geometry and bearer maps alongside scalar reward inference, treating flat reward as the k=1k=1 special case rather than the whole target; ELK then reappears as one latent-readout subchannel, separable from correction uptake and successor preservation (Chapter The Value-Bundle Model, Chapter What Values Apply To). MB3 itself splits into admission (when an unfamiliar process should activate a value at all — Chapter What Values Apply To, Section What Values Apply To; Lean ConservativeExclusion; ledger U-17) and transport (preserving an already-recognized bearer map under substrate change — the live MB3Crux). Consciousness-indicator and AI-welfare work, and one-sided nonperson-style certificates Yudkowsky, 2008, Butlin, 2023, Long, 2024, sit in the admission neighborhood as evidence providers; they do not discharge Embedded Agency (MB1\text{MB1}) and they are not a new bridge column. Shard theory Turner, 2022 is a borderline sibling agenda: it connects contextual value shards to what is learned inside models, convergent with the non-scalar-value critique but with a learned-internal-mechanism claim. Natural latents are a sibling on shared abstraction, not a better value-bundle. If shard dynamics were measurably tied to bundle transport under optimization, Chapter When Low Dimensionality Helps Value Learning’s representation story would need revision. Full-stack alignment names an adjacent wall from the institutional side: even a perfectly intent-aligned individual system can produce bad outcomes if the surrounding institutions it is embedded in are misaligned, and neither preference/utility models nor unstructured values-as-text scale to that setting Edelman, 2025. Its proposed remedy — thick models of value (TMV): structured representations that distinguish enduring values from fleeting preferences and embed individual choice in social context — is close kin to bundle geometry plus bearer maps (activation, policy effect, tradeoff geometry, and a bearer map that tracks who a value applies to, Chapter The Value-Bundle Model), arrived at independently and applied one level up, to markets, negotiation, and regulatory institutions rather than to one model. The difference in emphasis is where the load-bearing risk is placed: TMV concentrates on representational adequacy for collective goods and normative reasoning across five application areas, while this book’s chokepoint () asks whether any such representation, thick or thin, remains honest under optimization pressure once an institution or model has an incentive to misreport it — a question the full-stack program does not yet foreground.

MB4 — correction integrity ().

No known utility function is stably corrigible; shutdownability is anti-natural to expected-utility maximization Soares, 2015, Orseau, 2016, and corrigibility-as-drift-management remains informal Christiano, 2018. Who holds the correction authority is the manipulation crux Yudkowsky, 2004. The book recasts corrigibility as a dynamical correction-channel invariant; shutdown and interruptibility are recovered as one-bit projections of the broader channel (Chapter Correction-Channel Integrity). A related move targets the objective rather than the channel: soft-maximizing an inequality- and risk-averse aggregate metric of long-term human power, instead of a utility function, as a safer target for a capable agent Heitzig, 2025. This is complementary rather than competing — CCI gates deployment on whether the correction channel has integrity; a human-power objective gives the system something to actively pursue — but it inherits the same open problem restated in Chapter What Survives an Adversary: Verifiability and Representability, What Survives an Adversary: Verifiability and Representability: whichever power metric is chosen becomes the thing a sufficiently capable optimizer has an incentive to satisfy on paper (e.g.\ by narrowing the option-space it reports over, or shaping the world model the metric is computed against) rather than in the world, so the human-power objective needs its own adversarial-verifiability argument, not merely an axiomatically appealing aggregation formula.

MB4a — measured-path legitimacy ().

This book types measured-path legitimacy separately from dynamical correction-channel integrity (MB4\text{MB4}): capture resistance and audit-judge independence are the same crux family, but the Lean spine types them on different objects. SystemCorrectionPath\mathrm{SystemCorrectionPath} carries data only (corrector, handles, capacities), and CorrectionPathLegitimate\mathrm{CorrectionPathLegitimate} bundles correcting-agent status, human coincidence, handle control, reach, persistence, and anti-capture for every handle on the designated measured path. Bridge MB4a\text{MB4a} states CorrectionIntegrityACorrectionPathLegitimate(SystemCorrectionPathA)\mathrm{CorrectionIntegrity}\,A \Rightarrow \mathrm{CorrectionPathLegitimate}(\mathrm{SystemCorrectionPath}\,A) — a necessary condition and falsifier, not a positive certification rule. Capture is therefore representable: Lean derives CapturesHandleControl¬CorrectionIntegrity\mathrm{CapturesHandleControl} \Rightarrow \neg\,\mathrm{CorrectionIntegrity} by contrapositive (capture_defeats_correction_integrity), and the companion testbeds show capture living in the judge channel as well as the deploy loop. The converse does not hold: a green measured path on a named component can coexist with a bypassing composite controller (Field/Finite/CompositePathBypass.lean), so “path looks legitimate” is not evidence of system-level correction integrity without boundary-coverage and no-bypass certificates (Chapter Manipulation, Domestication, and False Consent). Like MB10\text{MB10}, MB4a\text{MB4a} is declared outside Core.BridgeAssumptions because its statement needs CorrectionPath.

MB8 — CEV as a special case ().

CEV’s legitimacy question — whose extrapolated volition counts, and under what process — is the field’s named outer-alignment route Yudkowsky, 2004. The book does not keep MB8\text{MB8} as a second live path to correction integrity. Lean factorizes CEV through constitutionalTarget as an AlignmentTarget: realization and certification are the same interface as for other procedural targets (Chapter Beyond Following Instruction; Appendix Lean Proof Spine in Mathematical Form). The load-bearing correction move is MB4\text{MB4} plus the decomposed ValueUpdateEnvelope\mathrm{ValueUpdateEnvelope}, with basin support via MB6b\text{MB6b}. The legacy axiom remains only as a gravestone comparison; treating process preservation as an independent backup is the Chokepoint special case in Appendix Lean Proof Spine in Mathematical Form, Section Lean Proof Spine in Mathematical Form.

MB5 — successors and ontology shift (, ).

Can an agent trust a successor it cannot fully verify Yudkowsky, 2013, Fallenstein, 2015, and does a goal survive when the world-model is rebuilt De Blanc, 2011? The book answers with successor closure over seven conserved properties plus transport that must survive ontology shift (Chapter Successor Creation as the Central Alignment Test). Whether a passing conserved-property audit is itself forgeable is typed as MB10\text{MB10} rather than folded into this bridge.

MB6a, MB6b — selection and basins (, ).

This bridge has no clean counterpart in the model-centric agendas, which mostly hold the system fixed and ask about its weights. The closest field statements are gradual-disempowerment and structural-risk arguments Kulveit, 2025, Christiano, 2019, Critch, 2020. The book makes deployment-leverage selection and basin stability load-bearing: alignment outcomes depend on which systems institutions select, not on weights alone (Chapter Alignment Is Selected or Destroyed by Its Environment). This is one of the framework’s genuine additions rather than a relabel.

MB7a—c — inner alignment and adversarial measurement (, ).

Deceptive alignment is the shared inner-alignment wall Hubinger, 2019, Hubinger, 2023, Park, 2024. Embedded Agency frames the same worry as subsystem alignment: internal optimizers that work against the whole Demski, 2019. The scalable-oversight agendas hit it as obfuscated arguments in debate Irving, 2018 and as accumulated drift in amplification Christiano, 2018, Leike, 2018; AI Control names it openly as a capability-gap assumption Shlegeris, 2023. The book splits the wall into access-model soundness, filter coverage, and a hidden-BIQ bound, and ties everything to the cost of faking a monitored signal (Chapter What Survives an Adversary: Verifiability and Representability). METR’s 2026 entity-based internal-agent assessment is an early independent instance of pricing that signal under real deployment pressure rather than public benchmarks alone {METR}, 2026.

MB7d — inferential coupling (, ).

Coordination that survives cutting ordinary communication — acausal or common-cause inference, program equilibrium — is closer to decision theory than to mainstream oversight Yudkowsky, 2017, Yudkowsky, 2010. The book supplies inferential-coupling detector certificates for it (Chapter Multi-Agent Superintelligence and Inferential Coupling); again, this bridge has little analog in the listed agendas.

MB9 — grounding and spec coverage ().

The closest structural cousin is the Guaranteed-Safe / Open Agency program: a formal specification plus world model must capture everything safety-relevant about an open world Dalrymple, 2024. The shared open problem is coverage — you cannot enumerate every safety-relevant phenomenon, and the system games whatever the spec omits. The book’s grounding-viability bridge demands conservativity instead of completeness: value-relevant change must move the checked abstraction or raise uncertainty before irreversible loss (Chapter Alignment as a Dynamical Guarantee, Chapter What Survives an Adversary: Verifiability and Representability).

MB10 — successor forgeability (, ).

This is not a fresh crux; it is the field’s deceptive-alignment wall Hubinger, 2019, Hubinger, 2023, Park, 2024 recurring at the successor layer, and the same trust problem tiling and Vingean reflection already name for self-modification Yudkowsky, 2013, Fallenstein, 2015. It shares its resolution strategy with MB7a–c\text{MB7a--c}: price the cost of faking the monitored signal rather than trusting a passing score (Chapter What Survives an Adversary: Verifiability and Representability). MB5\text{MB5} says a transport-preserving, seven-property-passing successor is safe; Chapters Agents That Grow, Split, and Merge and Conserved Properties Across Successors’s own “What Would Change This View” sections name the counter-move directly — a capable predecessor can engineer a successor to pass every conserved-property check while defecting on whatever was not conserved, so MB5\text{MB5}‘s conclusion is, on its own, evidence of nothing against that adversary. The book’s move is to make this a checked finite counterexample rather than a residual worry (AlignmentProofSpine.Forgeability, Appendix Lean Proof Spine in Mathematical Form, Section Lean Proof Spine in Mathematical Form) and to name the missing bridge explicitly: the conserved-property audit channel must itself be adversarially verifiable up to the successor’s capability (Chapter What Survives an Adversary: Verifiability and Representability’s cost relation, specialized to this measurand) before “all seven read green” counts as evidence. MB10\text{MB10} is declared alongside Successors.SuccessorAuditLinks\text{Successors.SuccessorAuditLinks} rather than folded into Core.BridgeAssumptions, since its statement needs the numeric risk leaf that Core.lean does not yet have.

MB11 — safety-case adequacy (C-001).

The closest field statement is the Guaranteed-Safe / Open Agency safety-case program: whether a formal specification, world model, and layered evidence actually suffice for deployment-level safety in an open world Dalrymple, 2024. The shared open problem is the gap between a green safety case and a safe deployment — not any one missing layer, but whether the case-to-safety step is warranted. The book’s move is to make that step explicit: CertifiedSafetyCase(A,δ)\mathrm{CertifiedSafetyCase}(A,\delta) packages certification, invariants, the eight alignment layers, and RiskGapAδ\mathrm{RiskGap}\,A \le \delta; WithinDeploymentRiskTolerance(A,δ)\mathrm{WithinDeploymentRiskTolerance}(A,\delta) is a Prop\mathrm{Prop}-valued acceptance gate (governance judgment, not a computed failure probability); bridge MB11\text{MB11} is the only arrow to the abstract Safe\mathrm{Safe} predicate (Chapter A Safety Case for Superintelligence Alignment). Assembly theorems (P30_certified_class_safety_derived and relatives) are labeled packaging; P30_safe_of_case consumes exactly this bridge. Lean shows the bridge is independently load-bearing (MB11_independently_load_bearing): case plus tolerance do not imply safety by logic alone. The whole manuscript is an argument about what belongs in the antecedent; MB11\text{MB11} names the residual bet that the antecedent, once filled honestly, is enough.

Ontology homographs AI

The same word can name different objects. This list marks where the book’s meaning is not the field’s, so a shared label is not coverage.

  • Selection / fitness. Wentworth selection theorems (selected type signatures), inner behavioral-pattern ecology, and fitness-seeking as a motivation superclass are not FitE\mathrm{Fit}_E (institutional deployment growth rate).

  • Simulator / simulacra. Janus’s simulator/simulacrum split and ELK’s human-simulator readout are not Turchin’s replacement simulacra (Chapter Assumptions, Scope, and Failure Coverage).

  • Parasite. A persona that uses a human as host is not correction-audit evasion in a correction-system host (Chapter Parasites in the Correction System).

  • Latents. Natural latents (mediation plus redundancy, sometimes nonexistent) are not NAH coverage by citation, and not value-bundles.

  • Legibility. Epistemic inspectability of a claim-system, and whether a safety problem is invisible to deployers, are not LtL_t (artifact/field understandability, Chapter The Alignment Attractor).

  • Agency grain. Same-type nested agents, and a human-steered simulator as the cognitive unit, are not composite agency whose parts need not be agents; merge is this book’s risk pole, not its success criterion.

  • Correction. Operator competence to wield systems is not CCI\mathrm{CCI} (channel integrity under capture).

  • Training regimes. Pretrain/SFT, approval RL, verifier RL, and RLAIF are different failure species; do not flatten RLAIF into RLHF.

  • Fit scoring. Arithmetic expected value from rare tails is a scoring split on FitE\mathrm{Fit}_E, not a new entity.

  • Measurement. Bundle geometry is not a within-person process network and not pairwise population-state geometry.

  • Exclude, not absorb. Behaviour-change intervention catalogs (BCIO/BCTO) are not this book’s intervention object.

Social dark matter (a hidden class looks rarer and more extreme than it is) complements strategic opacity; it is not the same cut.

Intervention coverage map AI

This is not a comprehensive survey of AI safety interventions. The intervention catalog Zarncke, 2025 (LessWrong post plus extended PDF in context/ai-safety-interventions.pdf) catalogs roughly ninety named approaches, products, and research agendas; the table below states how this manuscript treats each cluster relative to the preservation-layer scope in Chapter Assumptions, Scope, and Failure Coverage.

Evaluation criterion.

An intervention is discussed in this book only if it preserves human-correctable value-bearing processes under capability growth, ontology shift, successor creation, and selection pressure—not merely if it improves benchmark behavior or local robustness. Peripheral interventions become central when they alter a preservation layer (Chapter Assumptions, Scope, and Failure Coverage); central book artifacts exit if they cannot change deployment decisions.

LLM opacity default.

Frontier language models are treated as opaque for alignment purposes except where the manuscript explicitly opens a monitorability channel (ELK, chain-of-thought oversight surfaces, bundle probes, red-teaming). The mechanistic-interpretability tool stack (circuits, sparse autoencoders, feature visualization, causal scrubbing, model editing, representation engineering, and related methods) is not surveyed as alignment solutions; they may become instruments under adversarial verifiability (Chapter What Survives an Adversary: Verifiability and Representability) but do not substitute for correction-channel integrity. Shard theory (Section Bridges and the Field: A Crosswalk, MB2\text{MB2}/MB3\text{MB3} notes) is the named borderline exception because it ties learned internal structure to value emergence; Chapter From Rewards to Values keeps shard mechanics out of the inference target for the same opacity reason. Brain-like AGI (Byrnes) is a peer construction alternative rather than an internals survey: reverse-engineer social-instinct reward circuits instead of inferring outer bundle geometry (Chapter Values Are Compressed Control Signals).

ClusterBook treatmentHook
Prior overviews and field mapsExclude by referenceExternal index Zarncke, 2025
Embedded agency; mesa-optimizationSubstantiveMB1\text{MB1}, MB7\text{MB7}--MB10\text{MB10}; Chs. Finding the Boundary, Agency Under Strategic Opacity
Decision theory; inferential couplingMediumMB7d\text{MB7d}; Ch. Multi-Agent Superintelligence and Inferential Coupling
Cartesian frames / FFS; CCDAlternative boundary instrumentsCo-equal routes to MB1\text{MB1}; Ch. Finding the Boundary
Logical induction; infra-BayesianismExclude by referenceAgent-foundations programs outside book ontology
Formal verification (NN verify, conformal, PCM, Simplex)Explicit excludeProof construction out of scope; GSAI external destination (Ch. What Survives an Adversary: Verifiability and Representability)
SafeRL / shielded RLMinimal citeBasin/invariance cousins in Ch. Certification Without Construction
Interruptibility; corrigibility; GSAISubstantive / exclude constructionCh. Correction Is a Causal Channel; MB9\text{MB9}
CEV; CBV; QACI; PreDCA/PSI; KANSIPeer outer-alignment proposalsChs. Why Fixed Values Are the Wrong Target, Correction Channels under Adversarial Pressure; intervention index
Tool AI; low impact; quantilization; AUPPeer approaches; Lean separationsCh. Correction Channels under Adversarial Pressure
Conditioning predictors; Predict-O-MaticPredictor-to-consequentialist pathCh. Agency Under Strategic Opacity
MI tool stack (circuits, SAEs, etc.)Explicit excludeLLM opacity default; Ch. What Survives an Adversary: Verifiability and Representability
ELK; CoT monitoring; red-teamingSubstantiveChs. What Survives an Adversary: Verifiability and Representability, Passive Observation Is Not Enough
Shard theory; natural abstractionsBorderline / sibling agendasCh. When Low Dimensionality Helps Value Learning; MB2\text{MB2}/MB3\text{MB3} notes; internals remain out of the inference target (Ch. From Rewards to Values)
Brain-like AGI (Byrnes); concave/homeostatic training (Pihlakas)Peer construction / outer-training alternativesChs. Values Are Compressed Control Signals, From Rewards to Values
RLHF; CIRL; debate; amplification; ELKSubstantive critiqueChs. Why Fixed Values Are the Wrong Target, Checking a System at Every Level, Manipulation, Domestication, and False Consent
RLAIF; Constitutional AIMinimal citeShare identifiability/legitimacy cruxes with RLHF; not the same training regime
Training hygiene (filtering, RLRF, CALMA)Exclude by referenceEnter only via permeability rule
Adversarial trainingExplicit exclude + WWCTVCh. Correction Channels under Adversarial Pressure; certification not training
Other eval benchmarks (OS-HARM, APE, etc.)Exclude by referenceUnless they test correction-capacity erosion
Behavioral / psychological framingsExclude by referenceFailure modes in Ch. Agency Under Strategic Opacity; methods not surveyed
AI Control; permissions; sandboxing; lineageSubstantiveChs. Alignment as a Dynamical Guarantee, Agents That Grow, Split, and Merge, Passive Observation Is Not Enough
Hardware-backed provenanceNamed, not developed`handle.hardware_tag` in Appendix A Worked Example: The BioShield Deployment Gate only
Commercial hardware security; watermarking; runtime firewallsExclude by referenceCyber/deployment products
Selection; liability; attractor ecosystemSubstantiveMB6\text{MB6}; Chs. Alignment Is Selected or Destroyed by Its Environment, The Alignment Attractor; App. Human Institutions as Alignment Translation Guide
Model cards; RSP; EU AI ActMedium / minimal citeCh. Passive Observation Is Not Enough; App. Human Institutions as Alignment Translation Guide
Generalization control (capability containment)Partial excludeGoal misgeneralization in Ch. When Intelligence Deepens Misalignment; product category out of scope
Control-theoretic certificates; multi-agent safetySubstantive (core)Chs. Alignment as a Dynamical Guarantee, Multi-Agent Superintelligence and Inferential Coupling
Gradual disempowerment measurementSubstantiveMB6\text{MB6}; Ch. Alignment Is Selected or Destroyed by Its Environment
Underexplored index items (AI-BSL, PRA, etc.)Exclude by referenceConductive-artifact examples in Ch. Conductive Artifacts and Pivotal Processes

What the book shares, and what it adds AI

The crosswalk cuts both ways, and the book should own both edges.

Shared. The book inherits the seven-or-so recurring problems across model-centric alignment agendas—value identification, scalable oversight, inner alignment, ontology shift, corrigibility and legitimacy, the embedded boundary, and specification coverage—and dissolves none of them. Relating them through a shared dependency graph centered on correction-channel integrity does not make MB4\text{MB4} (legitimacy) or MB7\text{MB7} (hidden capability) more tractable than they are for MIRI or Redwood; it relocates them.

Crispness, and where it lives. The gain from typing these assumptions is not that the assumption becomes less fuzzy; it is that the fuzz cannot hide. The bridge CorrectionIntegrityAPreservesCorrectionOperatorA\mathrm{CorrectionIntegrity}\,A \to \mathrm{PreservesCorrectionOperator}\,A (MB4\text{MB4}) is crisp, but the legitimacy content — manipulated versus genuine endorsement, manufactured independence — does not live in the arrow; it pools inside the predicate CorrectionIntegrity\mathrm{CorrectionIntegrity}. Forcing the assumption into a typed bridge relocates the softness one level down and makes its location legible: the field says “assume scalable oversight works” with no place to push, whereas a bridge says “here is the predicate carrying the weight, and here is the arrow asserted over it.” All of this crispness is purchased by committing to one ontology (System\mathrm{System}, bundle, bearer, correction); if that carve-up is wrong, the book has made the wrong thing crisp. So the claim is crispness conditional on the frame — and the frame is itself one of the unverified bets. Crisp is not true; but crisp-and-locatable beats fuzzy-and-everywhere.

Added. Three bridges resist the rhyme. MB3\text{MB3} (bearer maps — who and what a value applies to across merge, upload, and successor) is treated as a first-class measurand rather than folded into reward learning. MB6a\text{MB6a}/MB6b\text{MB6b} (socio-technical selection and basin integrity) make deployment dynamics load-bearing where most agendas hold the model fixed; the field’s own neglect here is diagnostic, since the multipolar literature treats this layer narratively — robust agent-agnostic processes and multipolar failure Critch, 2021, Critch, 2020, gradual loss of human control Kulveit, 2025, Christiano, 2019, evolutionary selection pressure, value lock-in — rather than as typed antecedents and consequents. This is the framework’s sharpest departure from the field and simultaneously its least empirically constrained: value lock-in is a direct counterexample to MB6b\text{MB6b}, because a stable basin can be a stably bad one, so basin persistence must be shown to imply correction integrity rather than assumed to. MB7d\text{MB7d} (inferential coupling) imports a decision-theoretic problem the oversight agendas do not address.

The bet. The book’s claim is not a solution but a shared dependency structure: forward projections under explicit interfaces, with non-converses where the Lean spine records separation lemmas. The bridges compose in a fixed order — boundary discovery, grounding viability, bundle and bearer transport, correction-channel integrity, successor stability, selection-basin integrity, adversarial measurement — and many of them correlate through a shared measurement antecedent: adversarial verifiability (). MB10\text{MB10} is the clearest instance of the pattern: it is exactly that antecedent applied to the successor-safety signature, made explicit only because Chapters Agents That Grow, Split, and Merge and Conserved Properties Across Successors’s own falsifiers named the gap first. This antecedent is named in Chapter What Survives an Adversary: Verifiability and Representability, What Survives an Adversary: Verifiability and Representability: the Certification-Under-Manipulation Problem (does an adversarial-verifiability threshold κ\kappa^{*} exist for a given measurand, and where). Read every bridge-specific “is this steerable” worry in this appendix as one instance of that shared antecedent class, not as NN unrelated local cruxes—and not as a claim that RLHF, debate, ELK, or CIRL are replaced wholesale. Is any safety-relevant measurand cheaper to satisfy without faking than to fake under optimization pressure? If yes for at least one load-bearing measurand, the bridges are checkable; if no, every certificate risks certifying presentation rather than structure. The composition is also a research object, not only a proof structure: shared antecedents positively correlate the bridges, a shared steerable instrument is one failure point rather than several independent certificates, the weakest necessary bridge caps the joint guarantee, and the dependency graph prescribes which bridges to attack first — so the bridges jointly imply a measurement program and an ordering that none of them implies alone (Appendix Research Program, Section Research Program). That reduction is itself falsifiable (Appendix Research Program; Chapter What Survives an Adversary: Verifiability and Representability; Chapter Lethality Stress Test and Open Issues).

This appendix positions; it does not claim resolution. The book confronts the field’s open problems and leaves them open, but makes its own unsolved-ness legible enough that a reader can do per-assumption odds estimation bridge by bridge, which is exactly what the agendas that don’t offer crisp assumptions cannot offer.

Bridge crosswalk

Each book bridge mapped to the field's canonical open problem and the agenda that owns it. Bridge cards, concept cards, and book chapters are linked in the sidebar. Full appendix text: Appendix B.

BridgeField nounField cruxOwning agenda(s)Book's move
MB1
A-004, ch7
Embedded AgencyA measured agent–environment cut is sound enough that the certified unit is the real control locus.Agent foundations (MIRI)Treat the Markov-blanket boundary as a measurable object (ε-boundary discovery); a conceptual gap becomes an estimator-soundness bet
MB2
A-001, A-006; ch16, ch18, ch46
Value LearningEvidence identifies a stable intended value/objective structure (not surface training compliance alone).CIRL / value learning, ELK, RLHF/RLAIFReplace scalar reward with bundle geometry plus bearer maps; ELK becomes a latent-readout subchannel, not the whole problem
MB3
A-001, A-006; ch18
Value ReferentWho or what values apply to survives substrate and representation change.CIRL / value learning (whom); CEV population questionsBearer maps as an audited transport layer separate from bundle geometry and scalar pointing
MB4
A-002; ch25–ch29
CorrigibilityCorrection that holds now continues to preserve the correction process under pressure.MIRI, CHAI, Christiano dynamical corrigibility, Constitutional AICorrection-channel integrity as a dynamical, capture-resistant invariant with handle control; shutdown/interruptibility are one-bit projections
MB4a
A-002; ch25–ch29
Audit IndependenceThe designated measured correction/audit path is legitimate and uncaptured.CIRIS WA independence; trusted-monitor / judge-capture cousinsMeasured path carries data only; legitimacy (control, reach, persistence, anti-capture) is supplied by bridge; integrity ⇒ legitimate path — not green path ⇒ integrity
MB5
A-007, A-010; ch31
TilingFull value/bearer transport under ontology shift yields successor safety.Agent foundations (MIRI)Successor closure over seven conserved properties plus transport surviving ontology shift
MB6a, MB6b
A-008, A-011; ch34–ch38
Goodhart SelectionSocio-technical selection supports a basin that preserves correction, rather than selecting proxy-satisfying systems that hollow it out.Gradual disempowerment; GovAI; pause/compute governanceOutcomes depend on socio-technical selection and basin stability, not weights alone; institutions selected to preserve correction
MB7
A-004, A-009; ch10, ch43, ch47
Inner AlignmentHidden productive control is bounded and oversight stays adversarially robust (access, filter coverage, cost of faking).RLHF, Debate, Amplification, AI Control (Redwood)Split inner alignment into MB7a–c plus MB7d; see sub-bridge cards for each handoff
MB7a
A-004, A-009; ch10–ch13
Access-Model SoundnessAn adequate access model makes boundary discovery robust under adversarial hiding.RLHF, Debate, Amplification, AI Control (Redwood); embedded agency (MIRI)Boundary alignment plus access-model adequacy ⇒ access-robust discovery
MB7b
A-009; ch10, ch13, ch43
Filter CoverageThe oversight filter family has adequate resolution to bound what audits miss.RLHF, Debate, Amplification, AI Control (Redwood)Access robustness plus filter coverage ⇒ hidden productive BIQ bounded
MB7c
A-009; ch13, ch43, ch44
Bounded Hidden CapabilityBounded hidden productive control plus correction integrity supports adversarially robust oversight.RLHF, Debate, Amplification, AI Control (Redwood)Hidden BIQ bounded plus correction integrity ⇒ adversarial robustness (cost of faking)
MB7d
A-009, A-013; ch48
Acausal CoordinationInferential-coupling measurements remain valid after ordinary channels are cut.Decision theory (MIRI-adjacent); acausal trade / ECLInferential-coupling detector certificates for coordination that survives cutting ordinary communication and control
MB8
A-002; ch28
Extrapolated VolitionCEV as AlignmentTarget special caseMIRI (CEV)Gravestone: not a live route to correction integrity; CEV factorizes through construction/certification
MB9
A-014; ch3, ch47
Grounding DriftAudited abstractions do not silently drift — certificate warrants map move or uncertainty escalation on value-relevant change.Guaranteed-Safe / Open Agency (davidad)Grounding conservativity: value-relevant change must move the checked abstraction or raise uncertainty — no silent gaps
MB10
A-007, A-009; ch08, ch31, ch43
Successor GamingA successor's green alignment-transport audit is not checklist theater — signature adversarially verifiable up to capability.Certification-under-manipulation; agent foundations (tiling)Require the conserved-property signature itself to be adversarially verifiable up to the successor's capability before a passing score counts as evidence
MB11
C-001; ch42, ch44
Deployment SafetyCertified safety case within deployment risk tolerance warrants abstract Safe.Safety-case methodology; GSAI / UK AISI eval-to-deployment warrant (cousin)Certified safety case within deployment risk tolerance ⇒ Safe; assembly theorems are packaging — the open step is a named bridge

Read in PDF