Field

AI safety and alignment

A map of research agendas and how they connect to the hard problems of the AI alignment field.

The field of AI safety and alignment research asks how to build advanced artificial intelligence such that it remains beneficial to humans as capabilities grow. It ask: how can we keep humans able to correct mistakes and reduce catastrophic or existential risk from runaway AI processes.

Wikipedia’s article on AI alignment:

In the field of artificial intelligence (AI), alignment aims to steer AI systems toward a person’s or group’s intended goals, preferences, or ethical principles. An AI system is considered aligned if it advances the intended objectives. A misaligned AI system pursues unintended objectives.

Alignment is a subfield of AI safety, alongside robustness, monitoring, and control of AI.

Norbert Wiener said in 1960:

If we use, to achieve our purposes, a mechanical agency with whose operation we cannot interfere effectively […] we had better be quite sure that the purpose put into the machine is the purpose which we really desire.

Eliezer Yudkowsky’s version of the complexity of value is:

Any simple goal you try to describe that is All We Need To Program Into AIs is almost certainly wrong.

The field is not one research program but a multitude of agendas: labs, nonprofits, academic groups, governance institutes, and research collaborations that share vocabulary while disagreeing on terminology, near-term priorities, and what would count as success.

A map of the field

Useful starting points:

This site adds an overview of who the major agendas are, the cruxes of the field, and what evidence each has published on the cruxes. The cruxes are presented them as bridges, named conditional handoffs between the cruxes and toward overall alignment. For example, Embedded Agency/MB1 asks whether an agent–environment boundary is sound enough to trust, i.e., the embedded-agency point of the missing clear cut between “the model” and “the optimizer”. None of these cruxes are solved, but different agendas have made progress to different degree on eaach.

What you will find here

  1. Agenda cards — one page per research or advocacy program. Browse all via the Field agenda badge.
  2. Coverage matrix — The main table of agenda ad bridges below. Column headers link to bridge cards; row headers link to agenda cards. Cell tags link to the evidence catalog below.
  3. Evidence catalog — citations backing the matrix.

For term disambiguation across agendas, see the inter-agenda glossary (manuscript App E is synced separately). For how this project maps bridges to field cruxes, see Appendix B.

Bridge dependency map

How this project's bridge assumptions (MB1–MB11) depend on one another, labeled in field nouns from the coverage matrix below.

This map shows logical and safety-case assembly dependencies among named bridge handoffs — not a proof that real systems satisfy them. Lean checks conditional structure when the bridges hold; field evidence, measurement, and deployment class still govern whether each handoff is reliable in practice.

Click a diamond to open its bridge card. Red edges: logical antecedents between bridges. Black edges: safety-case assembly into deployment safety (MB11). Graphviz source (.dot)

Hub card: Bridge assumptions · Formal spine: Lean proof spine

Field coverage

Tags: C conceptual, T theory, S simulation, P practical, D empirical (SW), E empirical (other), O other — each links to the evidence catalog below.

AgendaEmbedded AgencyMB1Value LearningMB2Value ReferentMB3CorrigibilityMB4Audit IndependenceMB4aTilingMB5Goodhart SelectionMB6Inner AlignmentMB7Acausal CoordinationMB7dExtrapolated VolitionMB8Grounding DriftMB9Successor GamingMB10Deployment SafetyMB11
MIRIC1,2C3T4,5T4,5C6,8, T7C1,9T10,124C131C6, T7
RedwoodC11C11E13, C11C11
CHAI / FAR.AIT14,16, C15T14T17T17E78
ChristianoC18C18C19C19C20C22,23, T21C22
ARCC18C18C18
GSAIT24T128
Anthropic / GoodfireE25,26, C27E25P28E25,31,134, E152,153, C29E13P28,135
Google DeepMindC133T21
Apollo / Truthful AIE148E148,149,35, E73, C150E148E148
METRE36E36,38E36E36
ResolutionO39T40O129
Neglected approachesO41, T42, C144T42C43C144, E145
OrthogonalC147T45T45
WentworthC46,47,48C49,51, T50C52
Kosoy / IB & LTAT118,154,155, C132T119,123,155, C156,157T123,154, C156,157C121,156C120,156T122T118,154,155, C132
CIRISP53,108, C103,104,106, D105C107, D54,55,56, D109,110,111, P108C107, D109,110,111P112, C113D114, C115,116C104,117C108,130, P108
GovAI / UK AISIE57,59,146, E62, P58,61,63P58,64, E62,136C60, P63P63
Pause clusterP65P66P65
CLRC67,68,151C70
AI FuturesO77
ConjectureC79
Safeguarded AIE139, T141P140E139, T142T137,141T137,142, P140
MAI + CIPC85P143C85P143
TSAS81,82, T80T83,84T83,86T87, S88, C89T125, S88T90,91, S92T93, S94, P95T96,97, S98, E99T100C89,131T101T102T127

Evidence catalog

Primary sources behind matrix tags. External links open in a new tab.

IDAgendaBridgeTypeEvidenceSource
1MIRIMB1, MB7CEmbedded Agency: no clean agent–environment cut; subsystem alignment bucketDemski & Garrabrant 2019
2MIRIMB1CAgent Foundations technical agenda (embedded agency, delegation, decision theory)Soares & Fallenstein 2015
3MIRIMB2CValue learning under training ambiguity and ontology changeSoares 2015
4MIRIMB4, MB4aTCorrigibility: no known utility function stably corrigibleSoares & Fallenstein 2015
5MIRIMB4, MB4aTSafely interruptible agents (formal interruptibility)Orseau & Armstrong 2016
6MIRIMB5, MB10CTiling agents for self-modifying AI (successor trust)Yudkowsky 2013
7MIRIMB5, MB10TVingean reflection (reasoning about smarter successors)Fallenstein 2015
8MIRIMB5COntological crises in agents' value systemsDe Blanc 2011
9MIRIMB7CPredict-O-Matic: predictors becoming consequentialistsDemski 2019
10MIRIMB7dTFunctional Decision TheoryYudkowsky & Soares 2017
11RedwoodMB4, MB4a, MB7, MB10CAI Control: safety under intentional subversion / capability-gap assumptionShlegeris et al. 2023
13RedwoodMB7, MB10EAlignment faking in LLMs under training/eval pressureGreenblatt et al. 2024
14CHAI / FAR.AIMB2, MB3TCooperative inverse reinforcement learning (assistance games)Hadfield-Menell et al. 2016
15CHAI / FAR.AIMB2CHuman Compatible control problem framingRussell 2019
16CHAI / FAR.AIMB2TAttainable utility preservation (conservative agency)Turner et al. 2019
17CHAI / FAR.AIMB4TOff-switch game (shutdown incentive structure)Hadfield-Menell et al. 2017
18Christiano / ARCMB2, MB3, MB7CELK: human simulator vs direct translatorChristiano, Cotra & Xu 2021
19ChristianoMB4CCorrigibility as drift management (informal dynamical framing)Christiano 2018
20ChristianoMB6CWhat failure looks like (gradual disempowerment narrative)Christiano 2019
21ChristianoMB7TAI safety via debateIrving, Christiano & Amodei 2018
22ChristianoMB7, MB10CAmplification / scalable oversight under optimizationChristiano et al. 2018
23ChristianoMB7CScalable agent oversight problem statementLeike et al. 2018
24GSAIMB9TGuaranteed Safe AI framework (spec + world model coverage wall)Dalrymple et al. 2024
25Anthropic / GoodfireMB2, MB3, MB7EConstitutional AI: principles-as-feedback / RLAIF stackBai et al. 2022
26Anthropic / GoodfireMB2ERLHF ceiling and misspecification under optimizationCasper et al. 2023
27Anthropic / GoodfireMB2CConcrete problems in AI safety (pointing / scalable oversight lineage)Amodei et al. 2016
28Anthropic / GoodfireMB6, MB11PResponsible Scaling Policy (capability thresholds, deployment gates)Anthropic RSP 2024
29Anthropic / GoodfireMB7CConditioning predictors / anthropic capture failure modeHubinger 2023
31Anthropic / GoodfireMB7EInternal agent monitoring (eval target for external red teams)METR red-team of Anthropic monitoring 2026
35Apollo / Truthful AIMB7EScheming-in-the-wild OSINT incident corpus (field-adjacent)CLTR 2026 report
36METRMB6, MB7, MB10EFrontier Risk Report: entity-based internal-agent assessmentMETR 2026
38METRMB7ERed-teaming frontier agent monitoring under deployment pressureRein 2026
39ResolutionMB1OAutomation-first alignment research strategyResolution launch essay
40ResolutionMB9TSingular learning / formal pipeline bet (Timaeus lineage)Murfet 2025 SLT position
41Neglected approachesMB2ONeglected-approaches portfolio strategy (AE Studio alignment agenda)AE Studio alignment agenda; LessWrong mirror; AE Studio Research
42Neglected approachesMB2, MB4THuman-power objective as outer target (Heitzig line)Heitzig & Potham 2025
43Neglected approachesMB6CAI Safety Interventions field index (cross-cuts agendas)Zarncke 2025
45OrthogonalMB2, MB4TQACI formal outer-alignment goal lineLeake & Persson 2023
46WentworthMB1CBoundaries as directed Markov blankets (utility-theoretic cut)Wentworth, Boundaries I
47WentworthMB1CAgent boundaries aren't Markov blankets (critique of naive blanket cuts)Wentworth 2022
48WentworthMB1CSelection theorems program (agent type signatures under selection)Wentworth 2021
49WentworthMB2CPointers problem: values as function of humansWentworth 2020
50WentworthMB2TNatural latents (formal shared-abstraction program)Wentworth & Lorell 2023
51WentworthMB2CShard theory (contextual value shards in trained models)Turner & Udell 2022
52WentworthMB5COntology identification / diamond maximizer problem framingAgent-like structure posts
53CIRISMB1PNamed-identity bet: Verify+Lens on certified occurrence vs composite controllerAccord (public text); Architecture; CIRISVerify
54CIRISMB4DCIRISAgent / Verify / Lens / Proxy shipped stackCIRIS architecture; CIRISAgent
55CIRISMB4DDeferral / Wise Authority / emergency shutdown proceduresAccord (public text); CIRISAgent overview
56CIRISMB4D50/50 smoke battery (pin parity, prohibitions, conscience, proxy, verify types)CIRISAgent safety tests; CIRISNodeCore safety CI loop
103CIRISMB1CNamed-identity bet: holds for federation admission, fails as alignment unitAccord Book IX; CIRISVerify federation identity
104CIRISMB1, MB9CNEW-04 / L-01: no poly-time complete compositional-deception detectorAccord Book IX §9.4
105CIRISMB1DCIRISVerify: hardware-rooted identity, transparency log, capability licensingCIRISVerify README
106CIRISMB1CKey task: Verify+Lens green / WA-blind composite counterexample charterAccord compositional limits (NEW-04); CIRISLens README
107CIRISMB4CMB4 challenge surface + probe battery (causal bite, capture theater, shutdown scope)CIRISAgent integrity compliance
108CIRISMB1, MB4PSub-ASI validated scope; ASI protocol aspirational until RC gates (CRE, red-team)Accord scope / disclaimers
109CIRISMB4DProhibition system + WiseBus tier gating (24 unit tests)CIRISAgent safety tests
110CIRISMB4DH3ERE conscience execution helpers (14 unit tests)CIRISAgent safety tests
111CIRISMB4DCIRISProxy fail-closed billing on service/network errorsCIRISProxy
112CIRISMB6PFederation Orthogonality gate + CEG attestations (partnership admission)Accord Book IX §8.1+; Architecture
113CIRISMB6CPartial selection story inside federation only (not market-wide MB6)Architecture / federation
114CIRISMB7DCIRISLens Coherence Ratchet + Capacity Score (external witness pipeline)CIRISLens README
115CIRISMB7CL-02 detector-cooperation bet; Lens = triage not verdictCIRISLens README; CIRIS architecture (Lens triage)
116CIRISMB7CVerify: authentic ≠ ethical; measurements not verdictsCIRISVerify README
117CIRISMB9CBook IX explicit non-claim: no perfect safety under adversarial superintelligenceAccord Book IX
57GovAI / UK AISIMB6EGradual disempowerment: systemic risk from incremental AI developmentKulveit et al. 2025
58GovAI / UK AISIMB6, MB7PCompute governance and frontier-policy research pipelineGovAI compute governance report
59GovAI / UK AISIMB6EInternational control of powerful technology (governance analogies)GovAI Baruch-plan paper
60GovAI / UK AISIMB9CInstitutional translation of safety specs (policy-facing coverage)GovAI publications
61GovAI / UK AISIMB6PUK AISI frontier model testing mandateUK AISI eval lessons (2024)
62GovAI / UK AISIMB6, MB7ECheating behaviour in frontier model evaluationsUK AISI 2026
63GovAI / UK AISIMB6, MB9, MB11PStandards and pre-deployment testing (UK + US CAISI cluster)US NIST AI
64GovAI / UK AISIMB7PGovernment-led frontier eval binding on deploymentUK AISI Frontier AI Trends Report
65Pause clusterMB4, MB8POff-switch / pause priority in advocacy platformsPauseAI policy proposal
66Pause clusterMB6PMoratorium and verified-slowdown campaignsFLI pause letter
67CLRMB6CARCHES: multipolar and cooperation failure taxonomyCritch & Krueger 2020
68CLRMB6CMultipolar failure modes under competitionChristiano 2019 (multipolar post)
70CLRMB7dCEvidential cooperation / acausal trade lineFDT 2017
73Apollo / Truthful AIMB7EAI deception survey (field synthesis)Park et al. 2024
77AI FuturesMB6OAI 2027 scenario (schedule shapes for governance stress tests only)AI 2027 scenario summary
78CHAI / FAR.AIMB7EScalable oversight via partitioned human supervision (FAR.Lab)Yin et al. 2025
79ConjectureMB7CCognitive emulation / controllable LLM framingConjecture CoEm proposal
80TSAMB1TMB1 typed bridge + ε-boundary discovery (Lean + ch07)Ch. 7 (companion); Lean spine
81TSAMB1SEmbedded / lab boundary-discovery testbeds (interventional handles)Embedded simulation findings; Lab simulation findings
82TSAMB1SUAD / agency-detect (unsupervised boundary discovery from dynamics)Unsupervised Agent Discovery; agency-detect repo
83TSAMB2, MB3TBundle geometry + bearer maps (ch16, ch18)Ch. 16 (companion); Ch. 18 (companion)
84TSAMB2TLean CIRL / IRL non-identifiability projectionsLean spine; Field modules (GitHub)
85MAI + CIPMB2, MB6CFull-Stack Alignment (thick values, institutional amplification)Edelman et al. 2025
86TSAMB3TBearer-map transport under optimization (MB3 bridge)App B bridge crosswalk (companion)
87TSAMB4TCorrection-channel integrity invariant (Lean + ch26)Ch. 26 (companion); Lean spine
88TSAMB4SToy/lab correction-channel and capture scenariosToy simulation findings; Goal-agent simulation findings
89TSAMB4, MB8CCEV-process convergence as secondary route (not assumed)App B bridge crosswalk (companion)
90TSAMB5TSuccessor closure over seven conserved propertiesCh. 31 (companion)
91TSAMB5TOntology-shift transport (A-007, A-010)App B bridge crosswalk (companion)
92TSAMB5SGrow/split/merge successor stress testsGraded-lab simulation findings
93TSAMB6TSelection environment + deployment leverage (ch34)Ch. 34 (companion)
94TSAMB6SSelection / basin scenarios in graded-lab lineGraded-lab simulation findings
95TSAMB6PInstitutional translation appendix (App C)App C (companion)
96TSAMB7THidden productive BIQ bound + adversarial verifiability (A-009, ch43)Ch. 43 (companion)
97TSAMB7TLean ELK/debate separations (readout ⇏ correction)Lean spine
98TSAMB7SStrategic opacity / hidden-capability lab scenariosLab simulation findings
99TSAMB7EHubinger deceptive-alignment taxonomy as field wall (ch44 cite)Hubinger et al. 2019
100TSAMB7dTInferential-coupling detector certificates (ch35)Ch. 35 (companion)
101TSAMB9TGrounding conservativity vs GSAI completeness (ch dynamical guarantee)App B bridge crosswalk (companion)
102TSAMB10TSuccessor forgeability counterexample + audit-channel bridgeLean spine; Forgeability.lean (GitHub)
118Kosoy / IB & LTAMB1, MB9TInfra-Bayesianism: imprecise probabilities for nonrealizability / model misspecInfra-Bayesianism sequence intro; LessWrong tag: infra-Bayesianism
119Kosoy / IB & LTAMB2TLearning-theoretic agenda for AI alignment (regret-style guarantees)Kosoy, LTA overview (2018); Kosoy, LTA status (2023)
120Kosoy / IB & LTAMB7CDaemons / inner optimizers in learning-theoretic alignment framingKosoy, Taming daemons (2018 LTA); Kosoy, LTA status (2023)
121Kosoy / IB & LTAMB5CRSI / self-improvement treated in LTA (cousin to tiling/Vingean walls)Kosoy, Recursive self-improvement (2018 LTA); Kosoy, LTA status (2023)
122Kosoy / IB & LTAMB7dTInfra-Bayesian decision theory / imprecise-probability agentsInfra-Bayesianism sequence intro; LessWrong tag: infra-Bayesianism
123Kosoy / IB & LTAMB2, MB3TPreDCA / Physicalist Superimitation: precursor-utility outer-alignment (pointer via causal precursors)PreDCA tag; Kosoy, PSI section (LTA status 2023); Kosoy, PreDCA shortform (2022)
124MIRI / GarrabrantMB5, MB7dTLogical induction (logical uncertainty under bounded reasoning)Garrabrant et al. 2017
125TSAMB4aTMB4a measured-path legitimacy; capture defeats correction integrityLean spine; Correction.lean (GitHub)
126TSAMB8TMB8 CEV-process convergence (Lean bridge; secondary route)Lean spine
127TSAMB11TMB11 safety-case adequacy: certified case + tolerance → `Safe`Ch. 42 (companion); Lean spine
128GSAIMB11TConstructivist safety case / formal deployment guarantee programDalrymple et al. 2024
129ResolutionMB11OAutomated alignment risks under fuzzy research tasksIrving et al. 2026
130CIRISMB11CSub-ASI validated scope vs aspirational ASI protocol (honest safety-case disclaimers)Accord scope / disclaimers
131MIRI / YudkowskyMB8CCoherent extrapolated volition (field source for MB8 cousin)Yudkowsky 2004 CEV
132Kosoy / IB & LTAMB1, MB9CModel-class misspec / grain-of-truth (ambient MB1/MB9; Lean `Nonrealizability.lean`)Infra-Bayesianism sequence; Appel & Kosoy 2025 (robust regret); field-claim plan
133Google DeepMindMB1CDiscovering Agents (causal agent discovery from system dynamics)Kenton et al. 2022
134Anthropic / GoodfireMB7ECausally faithful mechanistic interpretability under interventionLange et al. 2023; Georg Lange (Foresight grantee)
135Anthropic / Goodfire (lab)MB11PFrontier Model Forum risk taxonomy and capability thresholdsFMF risk thresholds report
136GovAI / UK AISIMB7EAlignment evaluation case study; evaluation-awareness under agentic scaffoldsUK AISI 2026 alignment eval case study
137Safeguarded AIMB9, MB11TARIA Safeguarded AI programme (world models, specs, proof certificates)ARIA Safeguarded AI
139Safeguarded AIMB1, MB7EMulti-agent architecture security and collective-agency topologyHagag et al. 2026
140Safeguarded AIMB4a, MB11PAARM open runtime for checking and recording agent actionsAARM specification
141Safeguarded AIMB1, MB9TContainment Verification (action boundary with model as oracle)Moon & Varshney 2026
142Safeguarded AIMB7, MB11TZero-knowledge attestation for AI safety verificationBerrang ZK AI security
143MAI + CIPMB4a, MB8PAlignment assemblies and collective constitutional AICIP Alignment Assemblies
144Neglected approachesMB2, MB7CBrain-like AGI safety (social cognition / homeostatic alignment hypotheses)Byrnes brain-like AGI sequence
145Neglected approachesMB7ESelf-other overlap (SOO) fine-tuning against deceptive behaviorCarauleanu et al. 2024
146GovAI / UK AISIMB6EEmpirical disempowerment patterns in real-world LLM usageSharma et al. 2026
147OrthogonalMB1CEmbedded agency formalism (agent-foundations community lineage)Demski & Garrabrant 2019
148Apollo / Truthful AIMB6, MB7, MB10, MB11EIn-context scheming capabilities in frontier models (pre-deployment eval suite)Meinke et al. 2024
149Apollo / Truthful AIMB7ESituational Awareness Dataset (SAD) benchmark for LLMsLaine et al. 2024
150Apollo / Truthful AIMB7COut-of-context reasoning as foundation for situational awarenessBerglund et al. 2023
151CLRMB6CCooperative AI research agenda (CAIF lineage)Dafoe et al. 2020
152Anthropic / GoodfireMB7ESparse autoencoder monosemantic features (circuits interpretability)Bricken et al. 2023
153Anthropic / GoodfireMB7EScaling monosemantic features to Claude 3 SonnetTempleton et al. 2024
154Kosoy / IB & LTAMB1, MB3, MB9TInfra-Bayesian physicalism: bridge transform locates agent computations in a bird's-eye physical ontologyKosoy, Infra-Bayesian Physicalism
155Kosoy / IB & LTAMB1, MB2, MB9TRegret bounds for robust (multivalued / infra-Bayesian) online decision making under weak realizabilityAppel & Kosoy 2025 (COLT)
156Kosoy / IB & LTAMB2, MB3, MB5, MB7CLTA status 2023 integrative map (infra-Bayesianism ≠ whole agenda; Physicalist Superimitation / PreDCA outer strand)Kosoy, LTA status (2023)
157Kosoy / IB & LTAMB2, MB3CCommunity distillation of PreDCA protocol (precursor detection, classification, assistance)Soto, PreDCA distilled (2022)