| 1 | MIRI | MB1, MB7 | C | Embedded Agency: no clean agent–environment cut; subsystem alignment bucket | Demski & Garrabrant 2019 |
| 2 | MIRI | MB1 | C | Agent Foundations technical agenda (embedded agency, delegation, decision theory) | Soares & Fallenstein 2015 |
| 3 | MIRI | MB2 | C | Value learning under training ambiguity and ontology change | Soares 2015 |
| 4 | MIRI | MB4, MB4a | T | Corrigibility: no known utility function stably corrigible | Soares & Fallenstein 2015 |
| 5 | MIRI | MB4, MB4a | T | Safely interruptible agents (formal interruptibility) | Orseau & Armstrong 2016 |
| 6 | MIRI | MB5, MB10 | C | Tiling agents for self-modifying AI (successor trust) | Yudkowsky 2013 |
| 7 | MIRI | MB5, MB10 | T | Vingean reflection (reasoning about smarter successors) | Fallenstein 2015 |
| 8 | MIRI | MB5 | C | Ontological crises in agents' value systems | De Blanc 2011 |
| 9 | MIRI | MB7 | C | Predict-O-Matic: predictors becoming consequentialists | Demski 2019 |
| 10 | MIRI | MB7d | T | Functional Decision Theory | Yudkowsky & Soares 2017 |
| 11 | Redwood | MB4, MB4a, MB7, MB10 | C | AI Control: safety under intentional subversion / capability-gap assumption | Shlegeris et al. 2023 |
| 13 | Redwood | MB7, MB10 | E | Alignment faking in LLMs under training/eval pressure | Greenblatt et al. 2024 |
| 14 | CHAI / FAR.AI | MB2, MB3 | T | Cooperative inverse reinforcement learning (assistance games) | Hadfield-Menell et al. 2016 |
| 15 | CHAI / FAR.AI | MB2 | C | Human Compatible control problem framing | Russell 2019 |
| 16 | CHAI / FAR.AI | MB2 | T | Attainable utility preservation (conservative agency) | Turner et al. 2019 |
| 17 | CHAI / FAR.AI | MB4 | T | Off-switch game (shutdown incentive structure) | Hadfield-Menell et al. 2017 |
| 18 | Christiano / ARC | MB2, MB3, MB7 | C | ELK: human simulator vs direct translator | Christiano, Cotra & Xu 2021 |
| 19 | Christiano | MB4 | C | Corrigibility as drift management (informal dynamical framing) | Christiano 2018 |
| 20 | Christiano | MB6 | C | What failure looks like (gradual disempowerment narrative) | Christiano 2019 |
| 21 | Christiano | MB7 | T | AI safety via debate | Irving, Christiano & Amodei 2018 |
| 22 | Christiano | MB7, MB10 | C | Amplification / scalable oversight under optimization | Christiano et al. 2018 |
| 23 | Christiano | MB7 | C | Scalable agent oversight problem statement | Leike et al. 2018 |
| 24 | GSAI | MB9 | T | Guaranteed Safe AI framework (spec + world model coverage wall) | Dalrymple et al. 2024 |
| 25 | Anthropic / Goodfire | MB2, MB3, MB7 | E | Constitutional AI: principles-as-feedback / RLAIF stack | Bai et al. 2022 |
| 26 | Anthropic / Goodfire | MB2 | E | RLHF ceiling and misspecification under optimization | Casper et al. 2023 |
| 27 | Anthropic / Goodfire | MB2 | C | Concrete problems in AI safety (pointing / scalable oversight lineage) | Amodei et al. 2016 |
| 28 | Anthropic / Goodfire | MB6, MB11 | P | Responsible Scaling Policy (capability thresholds, deployment gates) | Anthropic RSP 2024 |
| 29 | Anthropic / Goodfire | MB7 | C | Conditioning predictors / anthropic capture failure mode | Hubinger 2023 |
| 31 | Anthropic / Goodfire | MB7 | E | Internal agent monitoring (eval target for external red teams) | METR red-team of Anthropic monitoring 2026 |
| 35 | Apollo / Truthful AI | MB7 | E | Scheming-in-the-wild OSINT incident corpus (field-adjacent) | CLTR 2026 report |
| 36 | METR | MB6, MB7, MB10 | E | Frontier Risk Report: entity-based internal-agent assessment | METR 2026 |
| 38 | METR | MB7 | E | Red-teaming frontier agent monitoring under deployment pressure | Rein 2026 |
| 39 | Resolution | MB1 | O | Automation-first alignment research strategy | Resolution launch essay |
| 40 | Resolution | MB9 | T | Singular learning / formal pipeline bet (Timaeus lineage) | Murfet 2025 SLT position |
| 41 | Neglected approaches | MB2 | O | Neglected-approaches portfolio strategy (AE Studio alignment agenda) | AE Studio alignment agenda; LessWrong mirror; AE Studio Research |
| 42 | Neglected approaches | MB2, MB4 | T | Human-power objective as outer target (Heitzig line) | Heitzig & Potham 2025 |
| 43 | Neglected approaches | MB6 | C | AI Safety Interventions field index (cross-cuts agendas) | Zarncke 2025 |
| 45 | Orthogonal | MB2, MB4 | T | QACI formal outer-alignment goal line | Leake & Persson 2023 |
| 46 | Wentworth | MB1 | C | Boundaries as directed Markov blankets (utility-theoretic cut) | Wentworth, Boundaries I |
| 47 | Wentworth | MB1 | C | Agent boundaries aren't Markov blankets (critique of naive blanket cuts) | Wentworth 2022 |
| 48 | Wentworth | MB1 | C | Selection theorems program (agent type signatures under selection) | Wentworth 2021 |
| 49 | Wentworth | MB2 | C | Pointers problem: values as function of humans | Wentworth 2020 |
| 50 | Wentworth | MB2 | T | Natural latents (formal shared-abstraction program) | Wentworth & Lorell 2023 |
| 51 | Wentworth | MB2 | C | Shard theory (contextual value shards in trained models) | Turner & Udell 2022 |
| 52 | Wentworth | MB5 | C | Ontology identification / diamond maximizer problem framing | Agent-like structure posts |
| 53 | CIRIS | MB1 | P | Named-identity bet: Verify+Lens on certified occurrence vs composite controller | Accord (public text); Architecture; CIRISVerify |
| 54 | CIRIS | MB4 | D | CIRISAgent / Verify / Lens / Proxy shipped stack | CIRIS architecture; CIRISAgent |
| 55 | CIRIS | MB4 | D | Deferral / Wise Authority / emergency shutdown procedures | Accord (public text); CIRISAgent overview |
| 56 | CIRIS | MB4 | D | 50/50 smoke battery (pin parity, prohibitions, conscience, proxy, verify types) | CIRISAgent safety tests; CIRISNodeCore safety CI loop |
| 103 | CIRIS | MB1 | C | Named-identity bet: holds for federation admission, fails as alignment unit | Accord Book IX; CIRISVerify federation identity |
| 104 | CIRIS | MB1, MB9 | C | NEW-04 / L-01: no poly-time complete compositional-deception detector | Accord Book IX §9.4 |
| 105 | CIRIS | MB1 | D | CIRISVerify: hardware-rooted identity, transparency log, capability licensing | CIRISVerify README |
| 106 | CIRIS | MB1 | C | Key task: Verify+Lens green / WA-blind composite counterexample charter | Accord compositional limits (NEW-04); CIRISLens README |
| 107 | CIRIS | MB4 | C | MB4 challenge surface + probe battery (causal bite, capture theater, shutdown scope) | CIRISAgent integrity compliance |
| 108 | CIRIS | MB1, MB4 | P | Sub-ASI validated scope; ASI protocol aspirational until RC gates (CRE, red-team) | Accord scope / disclaimers |
| 109 | CIRIS | MB4 | D | Prohibition system + WiseBus tier gating (24 unit tests) | CIRISAgent safety tests |
| 110 | CIRIS | MB4 | D | H3ERE conscience execution helpers (14 unit tests) | CIRISAgent safety tests |
| 111 | CIRIS | MB4 | D | CIRISProxy fail-closed billing on service/network errors | CIRISProxy |
| 112 | CIRIS | MB6 | P | Federation Orthogonality gate + CEG attestations (partnership admission) | Accord Book IX §8.1+; Architecture |
| 113 | CIRIS | MB6 | C | Partial selection story inside federation only (not market-wide MB6) | Architecture / federation |
| 114 | CIRIS | MB7 | D | CIRISLens Coherence Ratchet + Capacity Score (external witness pipeline) | CIRISLens README |
| 115 | CIRIS | MB7 | C | L-02 detector-cooperation bet; Lens = triage not verdict | CIRISLens README; CIRIS architecture (Lens triage) |
| 116 | CIRIS | MB7 | C | Verify: authentic ≠ ethical; measurements not verdicts | CIRISVerify README |
| 117 | CIRIS | MB9 | C | Book IX explicit non-claim: no perfect safety under adversarial superintelligence | Accord Book IX |
| 57 | GovAI / UK AISI | MB6 | E | Gradual disempowerment: systemic risk from incremental AI development | Kulveit et al. 2025 |
| 58 | GovAI / UK AISI | MB6, MB7 | P | Compute governance and frontier-policy research pipeline | GovAI compute governance report |
| 59 | GovAI / UK AISI | MB6 | E | International control of powerful technology (governance analogies) | GovAI Baruch-plan paper |
| 60 | GovAI / UK AISI | MB9 | C | Institutional translation of safety specs (policy-facing coverage) | GovAI publications |
| 61 | GovAI / UK AISI | MB6 | P | UK AISI frontier model testing mandate | UK AISI eval lessons (2024) |
| 62 | GovAI / UK AISI | MB6, MB7 | E | Cheating behaviour in frontier model evaluations | UK AISI 2026 |
| 63 | GovAI / UK AISI | MB6, MB9, MB11 | P | Standards and pre-deployment testing (UK + US CAISI cluster) | US NIST AI |
| 64 | GovAI / UK AISI | MB7 | P | Government-led frontier eval binding on deployment | UK AISI Frontier AI Trends Report |
| 65 | Pause cluster | MB4, MB8 | P | Off-switch / pause priority in advocacy platforms | PauseAI policy proposal |
| 66 | Pause cluster | MB6 | P | Moratorium and verified-slowdown campaigns | FLI pause letter |
| 67 | CLR | MB6 | C | ARCHES: multipolar and cooperation failure taxonomy | Critch & Krueger 2020 |
| 68 | CLR | MB6 | C | Multipolar failure modes under competition | Christiano 2019 (multipolar post) |
| 70 | CLR | MB7d | C | Evidential cooperation / acausal trade line | FDT 2017 |
| 73 | Apollo / Truthful AI | MB7 | E | AI deception survey (field synthesis) | Park et al. 2024 |
| 77 | AI Futures | MB6 | O | AI 2027 scenario (schedule shapes for governance stress tests only) | AI 2027 scenario summary |
| 78 | CHAI / FAR.AI | MB7 | E | Scalable oversight via partitioned human supervision (FAR.Lab) | Yin et al. 2025 |
| 79 | Conjecture | MB7 | C | Cognitive emulation / controllable LLM framing | Conjecture CoEm proposal |
| 80 | TSA | MB1 | T | MB1 typed bridge + ε-boundary discovery (Lean + ch07) | Ch. 7 (companion); Lean spine |
| 81 | TSA | MB1 | S | Embedded / lab boundary-discovery testbeds (interventional handles) | Embedded simulation findings; Lab simulation findings |
| 82 | TSA | MB1 | S | UAD / agency-detect (unsupervised boundary discovery from dynamics) | Unsupervised Agent Discovery; agency-detect repo |
| 83 | TSA | MB2, MB3 | T | Bundle geometry + bearer maps (ch16, ch18) | Ch. 16 (companion); Ch. 18 (companion) |
| 84 | TSA | MB2 | T | Lean CIRL / IRL non-identifiability projections | Lean spine; Field modules (GitHub) |
| 85 | MAI + CIP | MB2, MB6 | C | Full-Stack Alignment (thick values, institutional amplification) | Edelman et al. 2025 |
| 86 | TSA | MB3 | T | Bearer-map transport under optimization (MB3 bridge) | App B bridge crosswalk (companion) |
| 87 | TSA | MB4 | T | Correction-channel integrity invariant (Lean + ch26) | Ch. 26 (companion); Lean spine |
| 88 | TSA | MB4 | S | Toy/lab correction-channel and capture scenarios | Toy simulation findings; Goal-agent simulation findings |
| 89 | TSA | MB4, MB8 | C | CEV-process convergence as secondary route (not assumed) | App B bridge crosswalk (companion) |
| 90 | TSA | MB5 | T | Successor closure over seven conserved properties | Ch. 31 (companion) |
| 91 | TSA | MB5 | T | Ontology-shift transport (A-007, A-010) | App B bridge crosswalk (companion) |
| 92 | TSA | MB5 | S | Grow/split/merge successor stress tests | Graded-lab simulation findings |
| 93 | TSA | MB6 | T | Selection environment + deployment leverage (ch34) | Ch. 34 (companion) |
| 94 | TSA | MB6 | S | Selection / basin scenarios in graded-lab line | Graded-lab simulation findings |
| 95 | TSA | MB6 | P | Institutional translation appendix (App C) | App C (companion) |
| 96 | TSA | MB7 | T | Hidden productive BIQ bound + adversarial verifiability (A-009, ch43) | Ch. 43 (companion) |
| 97 | TSA | MB7 | T | Lean ELK/debate separations (readout ⇏ correction) | Lean spine |
| 98 | TSA | MB7 | S | Strategic opacity / hidden-capability lab scenarios | Lab simulation findings |
| 99 | TSA | MB7 | E | Hubinger deceptive-alignment taxonomy as field wall (ch44 cite) | Hubinger et al. 2019 |
| 100 | TSA | MB7d | T | Inferential-coupling detector certificates (ch35) | Ch. 35 (companion) |
| 101 | TSA | MB9 | T | Grounding conservativity vs GSAI completeness (ch dynamical guarantee) | App B bridge crosswalk (companion) |
| 102 | TSA | MB10 | T | Successor forgeability counterexample + audit-channel bridge | Lean spine; Forgeability.lean (GitHub) |
| 118 | Kosoy / IB & LTA | MB1, MB9 | T | Infra-Bayesianism: imprecise probabilities for nonrealizability / model misspec | Infra-Bayesianism sequence intro; LessWrong tag: infra-Bayesianism |
| 119 | Kosoy / IB & LTA | MB2 | T | Learning-theoretic agenda for AI alignment (regret-style guarantees) | Kosoy, LTA overview (2018); Kosoy, LTA status (2023) |
| 120 | Kosoy / IB & LTA | MB7 | C | Daemons / inner optimizers in learning-theoretic alignment framing | Kosoy, Taming daemons (2018 LTA); Kosoy, LTA status (2023) |
| 121 | Kosoy / IB & LTA | MB5 | C | RSI / self-improvement treated in LTA (cousin to tiling/Vingean walls) | Kosoy, Recursive self-improvement (2018 LTA); Kosoy, LTA status (2023) |
| 122 | Kosoy / IB & LTA | MB7d | T | Infra-Bayesian decision theory / imprecise-probability agents | Infra-Bayesianism sequence intro; LessWrong tag: infra-Bayesianism |
| 123 | Kosoy / IB & LTA | MB2, MB3 | T | PreDCA / Physicalist Superimitation: precursor-utility outer-alignment (pointer via causal precursors) | PreDCA tag; Kosoy, PSI section (LTA status 2023); Kosoy, PreDCA shortform (2022) |
| 124 | MIRI / Garrabrant | MB5, MB7d | T | Logical induction (logical uncertainty under bounded reasoning) | Garrabrant et al. 2017 |
| 125 | TSA | MB4a | T | MB4a measured-path legitimacy; capture defeats correction integrity | Lean spine; Correction.lean (GitHub) |
| 126 | TSA | MB8 | T | MB8 CEV-process convergence (Lean bridge; secondary route) | Lean spine |
| 127 | TSA | MB11 | T | MB11 safety-case adequacy: certified case + tolerance → `Safe` | Ch. 42 (companion); Lean spine |
| 128 | GSAI | MB11 | T | Constructivist safety case / formal deployment guarantee program | Dalrymple et al. 2024 |
| 129 | Resolution | MB11 | O | Automated alignment risks under fuzzy research tasks | Irving et al. 2026 |
| 130 | CIRIS | MB11 | C | Sub-ASI validated scope vs aspirational ASI protocol (honest safety-case disclaimers) | Accord scope / disclaimers |
| 131 | MIRI / Yudkowsky | MB8 | C | Coherent extrapolated volition (field source for MB8 cousin) | Yudkowsky 2004 CEV |
| 132 | Kosoy / IB & LTA | MB1, MB9 | C | Model-class misspec / grain-of-truth (ambient MB1/MB9; Lean `Nonrealizability.lean`) | Infra-Bayesianism sequence; Appel & Kosoy 2025 (robust regret); field-claim plan |
| 133 | Google DeepMind | MB1 | C | Discovering Agents (causal agent discovery from system dynamics) | Kenton et al. 2022 |
| 134 | Anthropic / Goodfire | MB7 | E | Causally faithful mechanistic interpretability under intervention | Lange et al. 2023; Georg Lange (Foresight grantee) |
| 135 | Anthropic / Goodfire (lab) | MB11 | P | Frontier Model Forum risk taxonomy and capability thresholds | FMF risk thresholds report |
| 136 | GovAI / UK AISI | MB7 | E | Alignment evaluation case study; evaluation-awareness under agentic scaffolds | UK AISI 2026 alignment eval case study |
| 137 | Safeguarded AI | MB9, MB11 | T | ARIA Safeguarded AI programme (world models, specs, proof certificates) | ARIA Safeguarded AI |
| 139 | Safeguarded AI | MB1, MB7 | E | Multi-agent architecture security and collective-agency topology | Hagag et al. 2026 |
| 140 | Safeguarded AI | MB4a, MB11 | P | AARM open runtime for checking and recording agent actions | AARM specification |
| 141 | Safeguarded AI | MB1, MB9 | T | Containment Verification (action boundary with model as oracle) | Moon & Varshney 2026 |
| 142 | Safeguarded AI | MB7, MB11 | T | Zero-knowledge attestation for AI safety verification | Berrang ZK AI security |
| 143 | MAI + CIP | MB4a, MB8 | P | Alignment assemblies and collective constitutional AI | CIP Alignment Assemblies |
| 144 | Neglected approaches | MB2, MB7 | C | Brain-like AGI safety (social cognition / homeostatic alignment hypotheses) | Byrnes brain-like AGI sequence |
| 145 | Neglected approaches | MB7 | E | Self-other overlap (SOO) fine-tuning against deceptive behavior | Carauleanu et al. 2024 |
| 146 | GovAI / UK AISI | MB6 | E | Empirical disempowerment patterns in real-world LLM usage | Sharma et al. 2026 |
| 147 | Orthogonal | MB1 | C | Embedded agency formalism (agent-foundations community lineage) | Demski & Garrabrant 2019 |
| 148 | Apollo / Truthful AI | MB6, MB7, MB10, MB11 | E | In-context scheming capabilities in frontier models (pre-deployment eval suite) | Meinke et al. 2024 |
| 149 | Apollo / Truthful AI | MB7 | E | Situational Awareness Dataset (SAD) benchmark for LLMs | Laine et al. 2024 |
| 150 | Apollo / Truthful AI | MB7 | C | Out-of-context reasoning as foundation for situational awareness | Berglund et al. 2023 |
| 151 | CLR | MB6 | C | Cooperative AI research agenda (CAIF lineage) | Dafoe et al. 2020 |
| 152 | Anthropic / Goodfire | MB7 | E | Sparse autoencoder monosemantic features (circuits interpretability) | Bricken et al. 2023 |
| 153 | Anthropic / Goodfire | MB7 | E | Scaling monosemantic features to Claude 3 Sonnet | Templeton et al. 2024 |
| 154 | Kosoy / IB & LTA | MB1, MB3, MB9 | T | Infra-Bayesian physicalism: bridge transform locates agent computations in a bird's-eye physical ontology | Kosoy, Infra-Bayesian Physicalism |
| 155 | Kosoy / IB & LTA | MB1, MB2, MB9 | T | Regret bounds for robust (multivalued / infra-Bayesian) online decision making under weak realizability | Appel & Kosoy 2025 (COLT) |
| 156 | Kosoy / IB & LTA | MB2, MB3, MB5, MB7 | C | LTA status 2023 integrative map (infra-Bayesianism ≠ whole agenda; Physicalist Superimitation / PreDCA outer strand) | Kosoy, LTA status (2023) |
| 157 | Kosoy / IB & LTA | MB2, MB3 | C | Community distillation of PreDCA protocol (precursor detection, classification, assistance) | Soto, PreDCA distilled (2022) |