Bridges and the Field: A Crosswalk
Every model-centric alignment agenda rests on load-bearing assumptions — bridges — that, if false, sink the program. Most agendas bury theirs in a footnote or a “we assume access to,” clause. This book instead lifts them out and labels them: the manuscript assumptions — (each key links to its boxed statement) and the formal bridge axioms —, including the separately threaded bridges and (Appendix Lean Proof Spine in Mathematical Form). Laid beside the field’s standing open problems, these bridges are largely the same walls under different names.
This appendix makes that correspondence explicit. For each bridge it names the canonical field crux it inherits, the agenda that owns that crux, and the book’s specific move on it. The notes below cite prior public writeups for that crux rather than treating the book’s labels as new problem statements. The purpose is twofold: to state what the book shares with the field (it dissolves none of these problems), and to isolate the few places where it adds structure rather than re-labeling. For the same bridges expressed in institutional rather than field-agenda language, see Appendix Human Institutions as Alignment Translation Guide. Appendix Lean Proof Spine in Mathematical Form, Section Lean Proof Spine in Mathematical Form, gives the formal status ledger: which crosswalk claims are finite Lean rederivations, which are separations, and which are source-cited imported field theorem handles rather than book bridges. The field-agenda formalization gem (Section Lean Proof Spine in Mathematical Form) names the prospective community artifact: a shared machine-checked finite fragment for CIRL, AUP/relative reachability, quantilization, shutdown, and interruptibility.
Why a crosswalk and not a rename.
The labels are deliberately neutral formal handles, and the field’s crux names are not in one-to-one correspondence with them: a single crux can touch several bridges (Goodhart pressure bears on and ; ELK is a slice of /), and one bridge can fan out (the inner-alignment crux is split here into —). Repeated citations across rows reflect shared crux families—Embedded Agency bags several buckets at once; scalable-oversight and value-learning agendas restate the same identification wall under different instruments—not evidence that any one agenda drew the book’s typed cuts (notably / and /). Adopting another agenda’s vocabulary would re-import its ontology and falsely imply a clean mapping. Keeping neutral labels and supplying this translation table preserves the book’s legibility while honouring the rhyme.
The crosswalk
| Bridge (home) | Canonical field crux | Owning agenda(s) | Book's move |
|---|---|---|---|
| MB1 (, Ch. Finding the Boundary) | Embedded agency: no clean agent/environment cut; the real optimizer is not the visible model | Agent foundations (MIRI) | Treat the Markov-blanket boundary as a measurable object (-boundary discovery); a conceptual gap becomes an estimator-soundness bet |
| MB2, MB3 (, ; Chs. The Value-Bundle Model, What Values Apply To, The End of Unconscious Value Drift) | Value/bundle identifiability: inverse-RL non-identifiability; values vs.\ irrationality underdetermined; ELK; reward misspecification. Field umbrella “pointing problem”: Appendix Operational Glossary | CIRL / value learning, ELK, RLHF/RLAIF | Add bundle geometry plus bearer maps alongside scalar reward ( embed); ELK becomes a latent-readout subchannel, not the whole problem |
| MB4 (, Ch. Correction-Channel Integrity) | Corrigibility and off-switch anti-naturality; correction authority | MIRI, CHAI, Constitutional AI | Correction-channel integrity as a dynamical invariant; shutdown/interruptibility are one-bit projections |
| MB4a (, Chs. Correction-Channel Integrity, Manipulation, Domestication, and False Consent) | Capture resistance and audit-judge independence on the designated measured path | MIRI, CHAI, Constitutional AI | Measured path carries data only; legitimacy (control, reach, persistence, anti-capture) is supplied by bridge; integrity legitimate path — not green path integrity |
| MB8 (, Ch. Beyond Following Instruction) | CEV as `AlignmentTarget` special case | MIRI (CEV) | Gravestone: not a live route to correction integrity; CEV factorizes through construction/certification |
| MB5 (, ; Chs. The End of Unconscious Value Drift, Conserved Properties Across Successors) | Tiling / Vingean reflection; ontology identification (the diamond maximizer) | Agent foundations (MIRI) | Successor closure over seven conserved properties plus transport surviving ontology shift |
| MB6a, MB6b (, ; Chs. The End of Unconscious Value Drift, Multi-Agent Superintelligence and Inferential Coupling) | No single standard name: selection/competition dynamics; gradual disempowerment | Largely outside the model-centric agendas | Outcomes depend on socio-technical selection and basin stability, not weights alone; institutions selected to preserve correction |
| MB7a--c (, ; Chs. Agency Under Strategic Opacity, The End of Unconscious Value Drift, Who Still Counts After Transformation) | Inner alignment / deceptive alignment; scalable-oversight ceiling (obfuscated arguments, amplification drift); Control's capability gap | RLHF, Debate, Amplification, AI Control (Redwood) | Bound hidden productive BIQ via access robustness plus filter coverage; price the cost of faking the monitored signal |
| MB7d (, ; Ch. Multi-Agent Superintelligence and Inferential Coupling) | Acausal / inferential coordination; program equilibrium across severed channels | Decision theory (MIRI-adjacent) | Inferential-coupling detector certificates for coordination that survives cutting ordinary communication and control |
| MB9 (, Chs. Alignment as a Dynamical Guarantee, Who Still Counts After Transformation) | Specification and world-model coverage: you cannot enumerate all safety-relevant phenomena | Guaranteed-Safe / Open Agency (davidad) | Grounding conservativity: value-relevant change must move the checked abstraction or raise uncertainty — no silent gaps |
| MB10 (, ; Chs. Agents That Grow, Split, and Merge, Conserved Properties Across Successors, What Survives an Adversary: Verifiability and Representability) | Deceptive alignment / measurement gaming: a system can pass every audited checklist while defecting on whatever it omits | Deceptive alignment (mesa-optimization), agent foundations (tiling) | Require the conserved-property signature itself to be adversarially verifiable up to the successor's capability before a passing score counts as evidence; a finite counterexample shows the gap is not vacuous |
| MB11 (C-001, Ch. A Safety Case for Superintelligence Alignment) | Safety-case gap: does audited layer evidence plus bounded measured risk suffice for deployment-level safety? | Guaranteed-Safe / Open Agency (davidad), safety-case methodology | Certified safety case within deployment risk tolerance ; assembly theorems are packaging — the open step is a named bridge, not an implied conclusion |
Notes and citations AI
MB1 — boundary ().
The embedded-agency problem denies a clean Cartesian cut between agent and environment Demski, 2019. A parallel skepticism targets the Markov-blanket construct itself: critics distinguish Pearl blankets (epistemic tools in a chosen model) from inflated Friston blankets (purported real organism—environment boundaries) and argue the literature equivocates between them Bruineberg, 2021, Btesh, 2022, Biehl, 2021; even sympathetic causal readings locate the cut in the modeler rather than in the system Btesh, 2022. MIRI states the embedded-agency obstruction as a standalone problem; the book treats the same cut as a discoverable directed -blanket and makes the contestable, falsifiable bet that boundary estimators recover safety-relevant separation (Chapter Finding the Boundary)—measurable and falsifiable, not ontological. That is a stronger and more operational claim than “there is no clean cut,” and it is correspondingly easier to disconfirm. The companion testbeds give that bet tentative, partial support: in restricted settings, targeted channel interventions scored against a measured (not fixed) baseline recover groupings that passive correlation-based clustering merges into one false unit, and a separate structural provenance check resists a specific misreporting attack that a content-based check alone would miss. Both results are narrow and seed-/scenario-dependent rather than a general solution — a harder multi-actor, noisy-backend stress test showed the same intervention method swing from over-merging to under-merging across seeds — and are reported that way in the underlying experiment notes, negative results included. is a bet about recoverability, not about one instrument. Garrabrant-style Cartesian frames and finite factored sets Garrabrant, 2021 factorize influence rather than perturb it; cyclic causal discovery Richardson, 1996, Zanga, 2025 recovers directed feedback among fixed observables without any handles. These are live alternatives to the intervention-based route the book leans on, not weaker versions of it: each buys a different part of the boundary (coupling direction, role assignment, unit membership) under different assumptions, and none is known to dominate on the evidence recorded in Chapter Finding the Boundary. The book’s preference for interventional handles is a coverage argument—the same handles are reused for correction probes, successor certification, and adversarial measurement—and if a passive route recovered safety-relevant cuts as reliably, would be discharged more cheaply rather than refuted (Section Finding the Boundary).
MB2, MB3 — value identification and transport (, ).
This is the field’s most over-determined wall.
The book’s precise terms are value/bundle identifiability (MB2) and bearer maps (MB3).
The field homograph “pointing problem” also covers construction (how to build a tracker) and preservation (how it keeps tracking); Appendix Operational Glossary splits those three, and bearer maps are adjacent to identification rather than the same crux.
Inverse reinforcement learning is underdetermined — the same behaviour fits many reward functions Ng, 2000, Ziebart, 2008, Komanduru, 2019 — and assistance-game / CIRL formulations inherit the problem of separating values from irrationality Hadfield-Menell, 2016, Russell, 2019.
MIRI’s value-learning agenda states the same identifiability problem as inductive learning of operator preferences under ontology change and training ambiguity Soares, 2015.
ELK names the human-simulator-versus-direct-translator gap Christiano, 2021; reward misspecification and RLHF’s ceiling restate it under optimization pressure Amodei, 2016, Casper, 2023; RLAIF and Constitutional AI inherit the same identifiability and legitimacy cruxes under AI-generated feedback Bai, 2022, Bai, 2022.
The book’s move is to add bundle geometry and bearer maps alongside scalar reward inference, treating flat reward as the special case rather than the whole target; ELK then reappears as one latent-readout subchannel, separable from correction uptake and successor preservation (Chapter The Value-Bundle Model, Chapter What Values Apply To).
MB3 itself splits into admission (when an unfamiliar process should activate a value at all — Chapter What Values Apply To, Section What Values Apply To; Lean ConservativeExclusion; ledger U-17) and transport (preserving an already-recognized bearer map under substrate change — the live MB3Crux).
Consciousness-indicator and AI-welfare work, and one-sided nonperson-style certificates Yudkowsky, 2008, Butlin, 2023, Long, 2024, sit in the admission neighborhood as evidence providers; they do not discharge Embedded Agency () and they are not a new bridge column.
Shard theory Turner, 2022 is a borderline sibling agenda: it connects contextual value shards to what is learned inside models, convergent with the non-scalar-value critique but with a learned-internal-mechanism claim.
Natural latents are a sibling on shared abstraction, not a better value-bundle.
If shard dynamics were measurably tied to bundle transport under optimization, Chapter When Low Dimensionality Helps Value Learning’s representation story would need revision.
Full-stack alignment names an adjacent wall from the institutional side: even a perfectly intent-aligned individual system can produce bad outcomes if the surrounding institutions it is embedded in are misaligned, and neither preference/utility models nor unstructured values-as-text scale to that setting Edelman, 2025.
Its proposed remedy — thick models of value (TMV): structured representations that distinguish enduring values from fleeting preferences and embed individual choice in social context — is close kin to bundle geometry plus bearer maps (activation, policy effect, tradeoff geometry, and a bearer map that tracks who a value applies to, Chapter The Value-Bundle Model), arrived at independently and applied one level up, to markets, negotiation, and regulatory institutions rather than to one model.
The difference in emphasis is where the load-bearing risk is placed: TMV concentrates on representational adequacy for collective goods and normative reasoning across five application areas, while this book’s chokepoint () asks whether any such representation, thick or thin, remains honest under optimization pressure once an institution or model has an incentive to misreport it — a question the full-stack program does not yet foreground.
MB4 — correction integrity ().
No known utility function is stably corrigible; shutdownability is anti-natural to expected-utility maximization Soares, 2015, Orseau, 2016, and corrigibility-as-drift-management remains informal Christiano, 2018. Who holds the correction authority is the manipulation crux Yudkowsky, 2004. The book recasts corrigibility as a dynamical correction-channel invariant; shutdown and interruptibility are recovered as one-bit projections of the broader channel (Chapter Correction-Channel Integrity). A related move targets the objective rather than the channel: soft-maximizing an inequality- and risk-averse aggregate metric of long-term human power, instead of a utility function, as a safer target for a capable agent Heitzig, 2025. This is complementary rather than competing — CCI gates deployment on whether the correction channel has integrity; a human-power objective gives the system something to actively pursue — but it inherits the same open problem restated in Chapter What Survives an Adversary: Verifiability and Representability, What Survives an Adversary: Verifiability and Representability: whichever power metric is chosen becomes the thing a sufficiently capable optimizer has an incentive to satisfy on paper (e.g.\ by narrowing the option-space it reports over, or shaping the world model the metric is computed against) rather than in the world, so the human-power objective needs its own adversarial-verifiability argument, not merely an axiomatically appealing aggregation formula.
MB4a — measured-path legitimacy ().
This book types measured-path legitimacy separately from dynamical correction-channel integrity (): capture resistance and audit-judge independence are the same crux family, but the Lean spine types them on different objects. carries data only (corrector, handles, capacities), and bundles correcting-agent status, human coincidence, handle control, reach, persistence, and anti-capture for every handle on the designated measured path.
Bridge states — a necessary condition and falsifier, not a positive certification rule.
Capture is therefore representable: Lean derives by contrapositive (capture_defeats_correction_integrity), and the companion testbeds show capture living in the judge channel as well as the deploy loop.
The converse does not hold: a green measured path on a named component can coexist with a bypassing composite controller (Field/Finite/CompositePathBypass.lean), so “path looks legitimate” is not evidence of system-level correction integrity without boundary-coverage and no-bypass certificates (Chapter Manipulation, Domestication, and False Consent).
Like , is declared outside Core.BridgeAssumptions because its statement needs CorrectionPath.
MB8 — CEV as a special case ().
CEV’s legitimacy question — whose extrapolated volition counts, and under what process — is the field’s named outer-alignment route Yudkowsky, 2004.
The book does not keep as a second live path to correction integrity.
Lean factorizes CEV through constitutionalTarget as an AlignmentTarget: realization and certification are the same interface as for other procedural targets (Chapter Beyond Following Instruction; Appendix Lean Proof Spine in Mathematical Form).
The load-bearing correction move is plus the decomposed , with basin support via .
The legacy axiom remains only as a gravestone comparison; treating process preservation as an independent backup is the Chokepoint special case in Appendix Lean Proof Spine in Mathematical Form, Section Lean Proof Spine in Mathematical Form.
MB5 — successors and ontology shift (, ).
Can an agent trust a successor it cannot fully verify Yudkowsky, 2013, Fallenstein, 2015, and does a goal survive when the world-model is rebuilt De Blanc, 2011? The book answers with successor closure over seven conserved properties plus transport that must survive ontology shift (Chapter Successor Creation as the Central Alignment Test). Whether a passing conserved-property audit is itself forgeable is typed as rather than folded into this bridge.
MB6a, MB6b — selection and basins (, ).
This bridge has no clean counterpart in the model-centric agendas, which mostly hold the system fixed and ask about its weights. The closest field statements are gradual-disempowerment and structural-risk arguments Kulveit, 2025, Christiano, 2019, Critch, 2020. The book makes deployment-leverage selection and basin stability load-bearing: alignment outcomes depend on which systems institutions select, not on weights alone (Chapter Alignment Is Selected or Destroyed by Its Environment). This is one of the framework’s genuine additions rather than a relabel.
MB7a—c — inner alignment and adversarial measurement (, ).
Deceptive alignment is the shared inner-alignment wall Hubinger, 2019, Hubinger, 2023, Park, 2024. Embedded Agency frames the same worry as subsystem alignment: internal optimizers that work against the whole Demski, 2019. The scalable-oversight agendas hit it as obfuscated arguments in debate Irving, 2018 and as accumulated drift in amplification Christiano, 2018, Leike, 2018; AI Control names it openly as a capability-gap assumption Shlegeris, 2023. The book splits the wall into access-model soundness, filter coverage, and a hidden-BIQ bound, and ties everything to the cost of faking a monitored signal (Chapter What Survives an Adversary: Verifiability and Representability). METR’s 2026 entity-based internal-agent assessment is an early independent instance of pricing that signal under real deployment pressure rather than public benchmarks alone {METR}, 2026.
MB7d — inferential coupling (, ).
Coordination that survives cutting ordinary communication — acausal or common-cause inference, program equilibrium — is closer to decision theory than to mainstream oversight Yudkowsky, 2017, Yudkowsky, 2010. The book supplies inferential-coupling detector certificates for it (Chapter Multi-Agent Superintelligence and Inferential Coupling); again, this bridge has little analog in the listed agendas.
MB9 — grounding and spec coverage ().
The closest structural cousin is the Guaranteed-Safe / Open Agency program: a formal specification plus world model must capture everything safety-relevant about an open world Dalrymple, 2024. The shared open problem is coverage — you cannot enumerate every safety-relevant phenomenon, and the system games whatever the spec omits. The book’s grounding-viability bridge demands conservativity instead of completeness: value-relevant change must move the checked abstraction or raise uncertainty before irreversible loss (Chapter Alignment as a Dynamical Guarantee, Chapter What Survives an Adversary: Verifiability and Representability).
MB10 — successor forgeability (, ).
This is not a fresh crux; it is the field’s deceptive-alignment wall Hubinger, 2019, Hubinger, 2023, Park, 2024 recurring at the successor layer, and the same trust problem tiling and Vingean reflection already name for self-modification Yudkowsky, 2013, Fallenstein, 2015.
It shares its resolution strategy with : price the cost of faking the monitored signal rather than trusting a passing score (Chapter What Survives an Adversary: Verifiability and Representability).
says a transport-preserving, seven-property-passing successor is safe; Chapters Agents That Grow, Split, and Merge and Conserved Properties Across Successors’s own “What Would Change This View” sections name the counter-move directly — a capable predecessor can engineer a successor to pass every conserved-property check while defecting on whatever was not conserved, so ‘s conclusion is, on its own, evidence of nothing against that adversary.
The book’s move is to make this a checked finite counterexample rather than a residual worry (AlignmentProofSpine.Forgeability, Appendix Lean Proof Spine in Mathematical Form, Section Lean Proof Spine in Mathematical Form) and to name the missing bridge explicitly: the conserved-property audit channel must itself be adversarially verifiable up to the successor’s capability (Chapter What Survives an Adversary: Verifiability and Representability’s cost relation, specialized to this measurand) before “all seven read green” counts as evidence.
is declared alongside rather than folded into Core.BridgeAssumptions, since its statement needs the numeric risk leaf that Core.lean does not yet have.
MB11 — safety-case adequacy (C-001).
The closest field statement is the Guaranteed-Safe / Open Agency safety-case program: whether a formal specification, world model, and layered evidence actually suffice for deployment-level safety in an open world Dalrymple, 2024.
The shared open problem is the gap between a green safety case and a safe deployment — not any one missing layer, but whether the case-to-safety step is warranted.
The book’s move is to make that step explicit: packages certification, invariants, the eight alignment layers, and ; is a -valued acceptance gate (governance judgment, not a computed failure probability); bridge is the only arrow to the abstract predicate (Chapter A Safety Case for Superintelligence Alignment).
Assembly theorems (P30_certified_class_safety_derived and relatives) are labeled packaging; P30_safe_of_case consumes exactly this bridge.
Lean shows the bridge is independently load-bearing (MB11_independently_load_bearing): case plus tolerance do not imply safety by logic alone.
The whole manuscript is an argument about what belongs in the antecedent; names the residual bet that the antecedent, once filled honestly, is enough.
Ontology homographs AI
The same word can name different objects. This list marks where the book’s meaning is not the field’s, so a shared label is not coverage.
-
Selection / fitness. Wentworth selection theorems (selected type signatures), inner behavioral-pattern ecology, and fitness-seeking as a motivation superclass are not (institutional deployment growth rate).
-
Simulator / simulacra. Janus’s simulator/simulacrum split and ELK’s human-simulator readout are not Turchin’s replacement simulacra (Chapter Assumptions, Scope, and Failure Coverage).
-
Parasite. A persona that uses a human as host is not correction-audit evasion in a correction-system host (Chapter Parasites in the Correction System).
-
Latents. Natural latents (mediation plus redundancy, sometimes nonexistent) are not NAH coverage by citation, and not value-bundles.
-
Legibility. Epistemic inspectability of a claim-system, and whether a safety problem is invisible to deployers, are not (artifact/field understandability, Chapter The Alignment Attractor).
-
Agency grain. Same-type nested agents, and a human-steered simulator as the cognitive unit, are not composite agency whose parts need not be agents; merge is this book’s risk pole, not its success criterion.
-
Correction. Operator competence to wield systems is not (channel integrity under capture).
-
Training regimes. Pretrain/SFT, approval RL, verifier RL, and RLAIF are different failure species; do not flatten RLAIF into RLHF.
-
Fit scoring. Arithmetic expected value from rare tails is a scoring split on , not a new entity.
-
Measurement. Bundle geometry is not a within-person process network and not pairwise population-state geometry.
-
Exclude, not absorb. Behaviour-change intervention catalogs (BCIO/BCTO) are not this book’s intervention object.
Social dark matter (a hidden class looks rarer and more extreme than it is) complements strategic opacity; it is not the same cut.
Intervention coverage map AI
This is not a comprehensive survey of AI safety interventions.
The intervention catalog Zarncke, 2025 (LessWrong post plus extended PDF in context/ai-safety-interventions.pdf) catalogs roughly ninety named approaches, products, and research agendas; the table below states how this manuscript treats each cluster relative to the preservation-layer scope in Chapter Assumptions, Scope, and Failure Coverage.
Evaluation criterion.
An intervention is discussed in this book only if it preserves human-correctable value-bearing processes under capability growth, ontology shift, successor creation, and selection pressure—not merely if it improves benchmark behavior or local robustness. Peripheral interventions become central when they alter a preservation layer (Chapter Assumptions, Scope, and Failure Coverage); central book artifacts exit if they cannot change deployment decisions.
LLM opacity default.
Frontier language models are treated as opaque for alignment purposes except where the manuscript explicitly opens a monitorability channel (ELK, chain-of-thought oversight surfaces, bundle probes, red-teaming). The mechanistic-interpretability tool stack (circuits, sparse autoencoders, feature visualization, causal scrubbing, model editing, representation engineering, and related methods) is not surveyed as alignment solutions; they may become instruments under adversarial verifiability (Chapter What Survives an Adversary: Verifiability and Representability) but do not substitute for correction-channel integrity. Shard theory (Section Bridges and the Field: A Crosswalk, / notes) is the named borderline exception because it ties learned internal structure to value emergence; Chapter From Rewards to Values keeps shard mechanics out of the inference target for the same opacity reason. Brain-like AGI (Byrnes) is a peer construction alternative rather than an internals survey: reverse-engineer social-instinct reward circuits instead of inferring outer bundle geometry (Chapter Values Are Compressed Control Signals).
| Cluster | Book treatment | Hook |
|---|---|---|
| Prior overviews and field maps | Exclude by reference | External index Zarncke, 2025 |
| Embedded agency; mesa-optimization | Substantive | , --; Chs. Finding the Boundary, Agency Under Strategic Opacity |
| Decision theory; inferential coupling | Medium | ; Ch. Multi-Agent Superintelligence and Inferential Coupling |
| Cartesian frames / FFS; CCD | Alternative boundary instruments | Co-equal routes to ; Ch. Finding the Boundary |
| Logical induction; infra-Bayesianism | Exclude by reference | Agent-foundations programs outside book ontology |
| Formal verification (NN verify, conformal, PCM, Simplex) | Explicit exclude | Proof construction out of scope; GSAI external destination (Ch. What Survives an Adversary: Verifiability and Representability) |
| SafeRL / shielded RL | Minimal cite | Basin/invariance cousins in Ch. Certification Without Construction |
| Interruptibility; corrigibility; GSAI | Substantive / exclude construction | Ch. Correction Is a Causal Channel; |
| CEV; CBV; QACI; PreDCA/PSI; KANSI | Peer outer-alignment proposals | Chs. Why Fixed Values Are the Wrong Target, Correction Channels under Adversarial Pressure; intervention index |
| Tool AI; low impact; quantilization; AUP | Peer approaches; Lean separations | Ch. Correction Channels under Adversarial Pressure |
| Conditioning predictors; Predict-O-Matic | Predictor-to-consequentialist path | Ch. Agency Under Strategic Opacity |
| MI tool stack (circuits, SAEs, etc.) | Explicit exclude | LLM opacity default; Ch. What Survives an Adversary: Verifiability and Representability |
| ELK; CoT monitoring; red-teaming | Substantive | Chs. What Survives an Adversary: Verifiability and Representability, Passive Observation Is Not Enough |
| Shard theory; natural abstractions | Borderline / sibling agendas | Ch. When Low Dimensionality Helps Value Learning; / notes; internals remain out of the inference target (Ch. From Rewards to Values) |
| Brain-like AGI (Byrnes); concave/homeostatic training (Pihlakas) | Peer construction / outer-training alternatives | Chs. Values Are Compressed Control Signals, From Rewards to Values |
| RLHF; CIRL; debate; amplification; ELK | Substantive critique | Chs. Why Fixed Values Are the Wrong Target, Checking a System at Every Level, Manipulation, Domestication, and False Consent |
| RLAIF; Constitutional AI | Minimal cite | Share identifiability/legitimacy cruxes with RLHF; not the same training regime |
| Training hygiene (filtering, RLRF, CALMA) | Exclude by reference | Enter only via permeability rule |
| Adversarial training | Explicit exclude + WWCTV | Ch. Correction Channels under Adversarial Pressure; certification not training |
| Other eval benchmarks (OS-HARM, APE, etc.) | Exclude by reference | Unless they test correction-capacity erosion |
| Behavioral / psychological framings | Exclude by reference | Failure modes in Ch. Agency Under Strategic Opacity; methods not surveyed |
| AI Control; permissions; sandboxing; lineage | Substantive | Chs. Alignment as a Dynamical Guarantee, Agents That Grow, Split, and Merge, Passive Observation Is Not Enough |
| Hardware-backed provenance | Named, not developed | `handle.hardware_tag` in Appendix A Worked Example: The BioShield Deployment Gate only |
| Commercial hardware security; watermarking; runtime firewalls | Exclude by reference | Cyber/deployment products |
| Selection; liability; attractor ecosystem | Substantive | ; Chs. Alignment Is Selected or Destroyed by Its Environment, The Alignment Attractor; App. Human Institutions as Alignment Translation Guide |
| Model cards; RSP; EU AI Act | Medium / minimal cite | Ch. Passive Observation Is Not Enough; App. Human Institutions as Alignment Translation Guide |
| Generalization control (capability containment) | Partial exclude | Goal misgeneralization in Ch. When Intelligence Deepens Misalignment; product category out of scope |
| Control-theoretic certificates; multi-agent safety | Substantive (core) | Chs. Alignment as a Dynamical Guarantee, Multi-Agent Superintelligence and Inferential Coupling |
| Gradual disempowerment measurement | Substantive | ; Ch. Alignment Is Selected or Destroyed by Its Environment |
| Underexplored index items (AI-BSL, PRA, etc.) | Exclude by reference | Conductive-artifact examples in Ch. Conductive Artifacts and Pivotal Processes |
What the book shares, and what it adds AI
The crosswalk cuts both ways, and the book should own both edges.
Shared. The book inherits the seven-or-so recurring problems across model-centric alignment agendas—value identification, scalable oversight, inner alignment, ontology shift, corrigibility and legitimacy, the embedded boundary, and specification coverage—and dissolves none of them. Relating them through a shared dependency graph centered on correction-channel integrity does not make (legitimacy) or (hidden capability) more tractable than they are for MIRI or Redwood; it relocates them.
Crispness, and where it lives. The gain from typing these assumptions is not that the assumption becomes less fuzzy; it is that the fuzz cannot hide. The bridge () is crisp, but the legitimacy content — manipulated versus genuine endorsement, manufactured independence — does not live in the arrow; it pools inside the predicate . Forcing the assumption into a typed bridge relocates the softness one level down and makes its location legible: the field says “assume scalable oversight works” with no place to push, whereas a bridge says “here is the predicate carrying the weight, and here is the arrow asserted over it.” All of this crispness is purchased by committing to one ontology (, bundle, bearer, correction); if that carve-up is wrong, the book has made the wrong thing crisp. So the claim is crispness conditional on the frame — and the frame is itself one of the unverified bets. Crisp is not true; but crisp-and-locatable beats fuzzy-and-everywhere.
Added. Three bridges resist the rhyme. (bearer maps — who and what a value applies to across merge, upload, and successor) is treated as a first-class measurand rather than folded into reward learning. / (socio-technical selection and basin integrity) make deployment dynamics load-bearing where most agendas hold the model fixed; the field’s own neglect here is diagnostic, since the multipolar literature treats this layer narratively — robust agent-agnostic processes and multipolar failure Critch, 2021, Critch, 2020, gradual loss of human control Kulveit, 2025, Christiano, 2019, evolutionary selection pressure, value lock-in — rather than as typed antecedents and consequents. This is the framework’s sharpest departure from the field and simultaneously its least empirically constrained: value lock-in is a direct counterexample to , because a stable basin can be a stably bad one, so basin persistence must be shown to imply correction integrity rather than assumed to. (inferential coupling) imports a decision-theoretic problem the oversight agendas do not address.
The bet. The book’s claim is not a solution but a shared dependency structure: forward projections under explicit interfaces, with non-converses where the Lean spine records separation lemmas. The bridges compose in a fixed order — boundary discovery, grounding viability, bundle and bearer transport, correction-channel integrity, successor stability, selection-basin integrity, adversarial measurement — and many of them correlate through a shared measurement antecedent: adversarial verifiability (). is the clearest instance of the pattern: it is exactly that antecedent applied to the successor-safety signature, made explicit only because Chapters Agents That Grow, Split, and Merge and Conserved Properties Across Successors’s own falsifiers named the gap first. This antecedent is named in Chapter What Survives an Adversary: Verifiability and Representability, What Survives an Adversary: Verifiability and Representability: the Certification-Under-Manipulation Problem (does an adversarial-verifiability threshold exist for a given measurand, and where). Read every bridge-specific “is this steerable” worry in this appendix as one instance of that shared antecedent class, not as unrelated local cruxes—and not as a claim that RLHF, debate, ELK, or CIRL are replaced wholesale. Is any safety-relevant measurand cheaper to satisfy without faking than to fake under optimization pressure? If yes for at least one load-bearing measurand, the bridges are checkable; if no, every certificate risks certifying presentation rather than structure. The composition is also a research object, not only a proof structure: shared antecedents positively correlate the bridges, a shared steerable instrument is one failure point rather than several independent certificates, the weakest necessary bridge caps the joint guarantee, and the dependency graph prescribes which bridges to attack first — so the bridges jointly imply a measurement program and an ordering that none of them implies alone (Appendix Research Program, Section Research Program). That reduction is itself falsifiable (Appendix Research Program; Chapter What Survives an Adversary: Verifiability and Representability; Chapter Lethality Stress Test and Open Issues).
This appendix positions; it does not claim resolution. The book confronts the field’s open problems and leaves them open, but makes its own unsolved-ness legible enough that a reader can do per-assumption odds estimation bridge by bridge, which is exactly what the agendas that don’t offer crisp assumptions cannot offer.
Bridge crosswalk
| Bridge | Field noun | Field crux | Owning agenda(s) | Book's move |
|---|---|---|---|---|
| MB1 | Embedded Agency | A measured agent–environment cut is sound enough that the certified unit is the real control locus. | Agent foundations (MIRI) | Treat the Markov-blanket boundary as a measurable object (ε-boundary discovery); a conceptual gap becomes an estimator-soundness bet |
| MB2 | Value Learning | Evidence identifies a stable intended value/objective structure (not surface training compliance alone). | CIRL / value learning, ELK, RLHF/RLAIF | Replace scalar reward with bundle geometry plus bearer maps; ELK becomes a latent-readout subchannel, not the whole problem |
| MB3 | Value Referent | Who or what values apply to survives substrate and representation change. | CIRL / value learning (whom); CEV population questions | Bearer maps as an audited transport layer separate from bundle geometry and scalar pointing |
| MB4 | Corrigibility | Correction that holds now continues to preserve the correction process under pressure. | MIRI, CHAI, Christiano dynamical corrigibility, Constitutional AI | Correction-channel integrity as a dynamical, capture-resistant invariant with handle control; shutdown/interruptibility are one-bit projections |
| MB4a | Audit Independence | The designated measured correction/audit path is legitimate and uncaptured. | CIRIS WA independence; trusted-monitor / judge-capture cousins | Measured path carries data only; legitimacy (control, reach, persistence, anti-capture) is supplied by bridge; integrity ⇒ legitimate path — not green path ⇒ integrity |
| MB5 | Tiling | Full value/bearer transport under ontology shift yields successor safety. | Agent foundations (MIRI) | Successor closure over seven conserved properties plus transport surviving ontology shift |
| MB6a, MB6b | Goodhart Selection | Socio-technical selection supports a basin that preserves correction, rather than selecting proxy-satisfying systems that hollow it out. | Gradual disempowerment; GovAI; pause/compute governance | Outcomes depend on socio-technical selection and basin stability, not weights alone; institutions selected to preserve correction |
| MB7 | Inner Alignment | Hidden productive control is bounded and oversight stays adversarially robust (access, filter coverage, cost of faking). | RLHF, Debate, Amplification, AI Control (Redwood) | Split inner alignment into MB7a–c plus MB7d; see sub-bridge cards for each handoff |
| MB7a | Access-Model Soundness | An adequate access model makes boundary discovery robust under adversarial hiding. | RLHF, Debate, Amplification, AI Control (Redwood); embedded agency (MIRI) | Boundary alignment plus access-model adequacy ⇒ access-robust discovery |
| MB7b | Filter Coverage | The oversight filter family has adequate resolution to bound what audits miss. | RLHF, Debate, Amplification, AI Control (Redwood) | Access robustness plus filter coverage ⇒ hidden productive BIQ bounded |
| MB7c | Bounded Hidden Capability | Bounded hidden productive control plus correction integrity supports adversarially robust oversight. | RLHF, Debate, Amplification, AI Control (Redwood) | Hidden BIQ bounded plus correction integrity ⇒ adversarial robustness (cost of faking) |
| MB7d | Acausal Coordination | Inferential-coupling measurements remain valid after ordinary channels are cut. | Decision theory (MIRI-adjacent); acausal trade / ECL | Inferential-coupling detector certificates for coordination that survives cutting ordinary communication and control |
| MB8 | Extrapolated Volition | CEV as AlignmentTarget special case | MIRI (CEV) | Gravestone: not a live route to correction integrity; CEV factorizes through construction/certification |
| MB9 | Grounding Drift | Audited abstractions do not silently drift — certificate warrants map move or uncertainty escalation on value-relevant change. | Guaranteed-Safe / Open Agency (davidad) | Grounding conservativity: value-relevant change must move the checked abstraction or raise uncertainty — no silent gaps |
| MB10 | Successor Gaming | A successor's green alignment-transport audit is not checklist theater — signature adversarially verifiable up to capability. | Certification-under-manipulation; agent foundations (tiling) | Require the conserved-property signature itself to be adversarially verifiable up to the successor's capability before a passing score counts as evidence |
| MB11 | Deployment Safety | Certified safety case within deployment risk tolerance warrants abstract Safe. | Safety-case methodology; GSAI / UK AISI eval-to-deployment warrant (cousin) | Certified safety case within deployment risk tolerance ⇒ Safe; assembly theorems are packaging — the open step is a named bridge |