A Map of the Field
External maps, surveys, and indexes for orienting to AI safety and alignment research.
Useful starting points for orienting to AI safety and alignment research:
- AISafety.com map — ecosystem map of organizations, programs, and resources (the clustering on this site rolls up map listings into coherent agenda rows).
- Mapping AI — U.S. policy-actor map (who can shape governance, what they believe, how they connect); insights are structural analyses of that database. Complementary to the research-ecosystem map above. Beliefs may be inferred; their “bridge builders” are connector people, not this site’s MB* cuts.
- AI Safety Interventions — index of roughly ninety named interventions across foundational theory, oversight, control, interpretability, governance, and underexplored routes; extended PDF.
- AI Alignment: A Comprehensive Survey (Ji et al., 2023) — academic survey of alignment problems and methods.
- Foundational challenges in assuring alignment and safety of LLMs (Ganguli et al., 2024) — Anthropic’s framing of eighteen foundational challenges.
- Open Problems in Technical AI Governance (Reuel et al., 2024) — policy and institutions adjacent to technical work.
- Center for AI Safety — field-building, statements, and course material.
- AI Alignment Forum / LessWrong — long-form research discourse and tag-based archives.
- Human Compatible (Russell, 2019) — assistance-games framing of beneficial AI under preference uncertainty.
- International AI Safety Report — periodic synthesis for policymakers (complementary to researcher-native maps).
- Textbook from the Future (Iliad) — theory megaproject and coordination instrument sketching a communal table of contents for foundational alignment; agenda card.
This site adds an overview of who the major agendas are, the cruxes of the field, and what evidence each has published on the cruxes. The cruxes are presented as bridge assumptions — named conditional handoffs between open problems and overall alignment. For example, Embedded Agency / MB1 asks whether an agent–environment boundary is sound enough to trust — the missing clear cut between “the model” and “the optimizer.” None of these cruxes are solved, but different agendas have made progress to different degrees on each.
For term disambiguation across agendas, see the inter-agenda glossary (manuscript App E is synced separately). For how this project maps bridges to field cruxes, see Appendix B — including Ontology homographs where the same English word names different objects. For the reverse question — what each agenda’s crux this map fails to represent — see What this map misses.