ARC (Alignment Research Center)
Can oversight elicit latent knowledge directly rather than through a human simulator, and will readout remain faithful under optimization (Inner Alignment)?
Introduction
ARC formalizes scalable oversight mechanisms, most prominently the ELK (Eliciting Latent Knowledge) problem, asking whether oversight can read out what a model knows rather than what it simulates.
Who carries it: Alignment Research Center (Geoffrey Irving pre-Resolution; ELK host)
What they aim to do. Formalize scalable oversight mechanisms, including latent knowledge elicitation, so that honest readout can be specified and tested.
The hard question. Can oversight elicit latent knowledge directly rather than through a human simulator, and will readout remain faithful under optimization (Inner Alignment)?
What they produce. ELK problem framing and formal oversight research hosted at the Alignment Research Center.
Key terms. Key terms include ELK, Eliciting Latent Knowledge, human simulator, and direct translator.
Related field cruxes. Value Learning; Value Referent; Inner Alignment
What they contribute. ELK as the canonical naming of the latent readout problem in scalable oversight.
How this project treats it. Successful readout does not imply correction uptake; this project treats correction-channel integrity as a separate requirement from eliciting latent knowledge, and adds Value Learning and Value Referent structure beyond readout alone.
Links
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.