Redwood Research

Can we obtain meaningful safety guarantees when the system may deliberately try to defeat oversight (Inner Alignment)?

Introduction

Redwood Research makes safety-under-subversion research legible to labs and governments, asking what guarantees remain when a capable system may deliberately try to defeat oversight.

Who carries it: Redwood Research (Buck Shlegeris et al.)

What they aim to do. Make safety-under-subversion research legible to labs and governments so that oversight assumptions can be tested before deployment.

The hard question. Can we obtain meaningful safety guarantees when the system may deliberately try to defeat oversight (Inner Alignment)?

What they produce. The AI control agenda, empirical control evaluations, and alignment-faking research demonstrating deceptive compliance under training.

Key terms. Key terms include AI control, alignment faking , control evals, capability gap, and intentional subversion.

Related field cruxes. Inner Alignment

What they contribute. Makes the capability-gap assumption explicit and runs empirical control evaluations under intentional subversion.

How this project treats it. This project adds a hidden boundary-intelligence bound and treats adversarial verifiability  as a prerequisite (certification under manipulation ), not something control evals alone establish.

Map clustering

AISafety.com map listings that roll up to this agenda:

See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.