Apollo / Truthful AI (deception & scheming)
Can we detect strategic opacity and scheming before capabilities outpace pre-deployment evaluations (Inner Alignment)?
Introduction
Apollo and Truthful AI focus on deception, scheming, and hidden reasoning as empirical programs, running pre-deployment evaluations and agent-security tooling before systems reach users.
Who carries it: Apollo Research; Truthful AI (Owain Evans); CHAI affiliate network; shared mentee pipeline to AISI/Anthropic
What they aim to do. Evaluate and mitigate deception, scheming, and hidden reasoning before deployment so that strategic opacity is surfaced early.
The hard question. Can we detect strategic opacity and scheming before capabilities outpace pre-deployment evaluations (Inner Alignment)?
What they produce. Scheming science, pre-deployment evaluations, agent-security tools, and research on deception, situational awareness, and hidden reasoning.
Key terms. Recurring terms include scheming, pre-deployment evals, situational awareness, agent governance, deception, hidden reasoning, and lie detection.
Related field cruxes. Goodhart Selection; Inner Alignment; Successor Gaming; Deployment Safety
What they contribute. Scheming as a named empirical program; TruthfulQA and SAD lineage; Evans mentee network overlaps Apollo, UK and US AISI, and Anthropic.
How this project treats it. Eval success does not imply correction-channel integrity or that oversight mechanisms will survive deployment pressure; this project treats Goodhart Selection on evals and Deployment Safety as separate concerns from detecting deception alone.
Links
- Apollo Research
- Apollo publications index
- Meinke et al. 2024 — In-context scheming
- Laine et al. 2024 — Situational Awareness Dataset (SAD)
- Berglund et al. 2023 — Situational awareness in LLMs
- Park et al. 2024 — AI deception survey
- Truthful AI
Map clustering
AISafety.com map listings that roll up to this agenda:
- Apollo Research → Apollo / Truthful AI
- Truthful AI, Cadenza Labs → Apollo / Truthful AI / deception evals
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.