Christiano lineage
Can oversight stay honest when arguments can be obfuscated, judge preferences drift over time, and latent readout may diverge from behavior (Inner Alignment)?
Introduction
Paul Christiano’s lineage developed scalable oversight protocols—debate, amplification, recursive reward modeling—and the ELK problem, aiming for oversight that stays honest under optimization pressure.
Who carries it: Paul Christiano, ARC-adjacent ELK, former OpenAI alignment
What they aim to do. Build scalable oversight mechanisms that remain honest under optimization pressure as systems grow more capable.
The hard question. Can oversight stay honest when arguments can be obfuscated, judge preferences drift over time, and latent readout may diverge from behavior (Inner Alignment)?
What they produce. Debate, amplification, and ELK posts; corrigibility notes emphasizing dynamical rather than one-shot properties.
Key terms. Signature concepts include debate, amplification, ELK, recursive reward modeling, corrigibility (understood dynamically), and scalable oversight.
Related field cruxes. Value Learning; Value Referent; Corrigibility; Inner Alignment
What they contribute. The scalable-oversight protocol family (debate, amplification) and the ELK problem statement as a canonical framing of latent readout.
How this project treats it. This project separates correction-channel integrity from readout: successful latent elicitation does not imply that corrections will be accepted (Corrigibility).
Links
Map clustering
AISafety.com map listings that roll up to this agenda:
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.