MIRI
There may be no clean cut between an AI and its environment (Embedded Agency); corrigibility may be anti-natural; and successor systems may not inherit trust under ontology change (Tiling).
Introduction
MIRI studies agent foundations—the theoretical problems that arise when building highly capable AI—and has increasingly emphasized policy advocacy for pause and off-switch measures when technical solutions look insufficient.
Who carries it: Machine Intelligence Research Institute
What they aim to do. Steer transformative AI away from extinction risk by framing alignment as an agent-foundations problem and prioritizing policy measures such as pause and off-switch when technical progress stalls.
The hard question. There may be no clean cut between an AI and its environment (Embedded Agency); corrigibility may be anti-natural; and successor systems may not inherit trust under ontology change (Tiling).
What they produce. The Agent Foundations technical agenda (including Embedded Agency work) and policy or pause outreach, notably the 2024 strategy update.
Key terms. Core terms include embedded agency, corrigibility, tiling, Vingean reflection, value learning, hard pause, and off-switch.
Related field cruxes. Embedded Agency; Corrigibility; Tiling; Extrapolated Volition
What they contribute. Canonical problem statements for embedded agency and corrigibility, plus a decision-theoretic lineage that shaped much of the field’s early vocabulary.
How this project treats it. This project treats boundary discovery as an operational bet on Embedded Agency, and distinguishes trajectory-level correction-channel integrity from a single-bit shutdown mechanism.
Links
Map clustering
AISafety.com map listings that roll up to this agenda:
- MIRI, AI StopWatch, Arbital (legacy canon) → MIRI
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.