Google DeepMind (safety)
Can oversight and safety research co-scale with capabilities, including under deceptive alignment and Inner Alignment risk—the same frontier-lab crux shared with other major labs?
Introduction
Google DeepMind integrates safety research into a large capabilities org, producing empirical alignment work and agent-discovery methods alongside Gemini-era deployment.
Who carries it: Google DeepMind safety / alignment researchers
What they aim to do. Build advanced AI with internal safety research integrated into the capabilities organization from the start.
The hard question. Can oversight and safety research co-scale with capabilities, including under deceptive alignment and Inner Alignment risk—the same frontier-lab crux shared with other major labs?
What they produce. Safety research blog output, Gemini-era alignment work, and the Discovering Agents program for causal agent discovery from system dynamics.
Key terms. Key terms include scalable alignment, safety research, evaluation, Gemini deployment, discovering agents, and causal agent discovery.
Related field cruxes. Embedded Agency; Inner Alignment
What they contribute. Large-scale empirical alignment research capacity and causal agent discovery methods that infer agents from system dynamics.
How this project treats it. Corporate safety teams do not replace an independent measurement spine for Embedded Agency and boundary discovery ; this project treats boundary and inner-alignment questions as only partially addressed by in-house teams.
Links
- Google DeepMind
- DeepMind Safety Research blog
- Discovering Agents
- Leike et al. 2018 — Scalable agent oversight
Map clustering
AISafety.com map listings that roll up to this agenda:
- Google DeepMind, DeepMind Safety Research → Google DeepMind safety
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.