Google DeepMind (safety)

Can oversight and safety research co-scale with capabilities, including under deceptive alignment and Inner Alignment risk—the same frontier-lab crux shared with other major labs?

Introduction

Google DeepMind integrates safety research into a large capabilities org, producing empirical alignment work and agent-discovery methods alongside Gemini-era deployment.

Who carries it: Google DeepMind safety / alignment researchers

What they aim to do. Build advanced AI with internal safety research integrated into the capabilities organization from the start.

The hard question. Can oversight and safety research co-scale with capabilities, including under deceptive alignment  and Inner Alignment risk—the same frontier-lab crux shared with other major labs?

What they produce. Safety research blog output, Gemini-era alignment work, and the Discovering Agents program for causal agent discovery from system dynamics.

Key terms. Key terms include scalable alignment, safety research, evaluation, Gemini deployment, discovering agents, and causal agent discovery.

Related field cruxes. Embedded Agency; Inner Alignment

What they contribute. Large-scale empirical alignment research capacity and causal agent discovery methods that infer agents from system dynamics.

How this project treats it. Corporate safety teams do not replace an independent measurement spine for Embedded Agency and boundary discovery ; this project treats boundary and inner-alignment questions as only partially addressed by in-house teams.

Map clustering

AISafety.com map listings that roll up to this agenda:

See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.