MIRI

There may be no clean cut between an AI and its environment (Embedded Agency); corrigibility may be anti-natural; and successor systems may not inherit trust under ontology change (Tiling).

Introduction

MIRI studies agent foundations—the theoretical problems that arise when building highly capable AI—and has increasingly emphasized policy advocacy for pause and off-switch measures when technical solutions look insufficient.

Who carries it: Machine Intelligence Research Institute

What they aim to do. Steer transformative AI away from extinction risk by framing alignment as an agent-foundations problem and prioritizing policy measures such as pause and off-switch when technical progress stalls.

The hard question. There may be no clean cut between an AI and its environment (Embedded Agency); corrigibility may be anti-natural; and successor systems may not inherit trust under ontology change (Tiling).

What they produce. The Agent Foundations technical agenda (including Embedded Agency work) and policy or pause outreach, notably the 2024 strategy update.

Key terms. Core terms include embedded agency, corrigibility, tiling, Vingean reflection, value learning, hard pause, and off-switch.

Related field cruxes. Embedded Agency; Corrigibility; Tiling; Extrapolated Volition

What they contribute. Canonical problem statements for embedded agency and corrigibility, plus a decision-theoretic lineage that shaped much of the field’s early vocabulary.

How this project treats it. This project treats boundary discovery  as an operational bet on Embedded Agency, and distinguishes trajectory-level correction-channel integrity  from a single-bit shutdown mechanism.

Map clustering

AISafety.com map listings that roll up to this agenda:

See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.