Anthropic / Goodfire (lab & MI stack)
Can scaling policies and interpretability keep pace with capability growth, including under strategic opacity and Goodhart Selection pressure?
Introduction
Anthropic builds frontier models under staged safety commitments (RSP), while adjacent teams and investments pursue the MI stack—mechanistic interpretability research, editable internal representations, and cross-lab tooling infrastructure (Goodfire, Transluce, Neuronpedia).
Who carries it: Anthropic PBC; Goodfire; Transluce; Neuronpedia (infra); Georg Lange (causal-faithfulness seam)
What they aim to do. Build capable systems with staged safety commitments, interpretability, and editable internal representations that can be audited and adjusted.
The hard question. Can scaling policies and interpretability keep pace with capability growth, including under strategic opacity and Goodhart Selection pressure?
What they produce. The Responsible Scaling Policy (RSP), Constitutional AI, interpretability research, Ember and editable representations, and FMF capability thresholds.
Key terms. Recurring terms include RSP, constitutional AI, mechanistic interpretability, circuits, features, steering, capability thresholds, and causal faithfulness.
Related field cruxes. Value Learning; Value Referent; Goodhart Selection; Inner Alignment; Successor Gaming; Deployment Safety
What they contribute. Industry RSP template, conditioning-predictor failure modes, Goodfire mechanistic-interpretability tooling (Anthropic investment; Apollo and DeepMind mechanistic interpretability lineage on team), and cross-lab Neuronpedia infrastructure.
How this project treats it. This project requires correction-channel integrity and adversarial verifiability ; a lab RSP is not the same as a preservation-layer certificate (Deployment Safety), and interpretability progress does not by itself resolve Successor Gaming or full Inner Alignment risk.
Links
- Anthropic
- Anthropic Research (index)
- Bricken et al. 2023 — Towards Monosemanticity
- Templeton et al. 2024 — Scaling Monosemanticity (Claude 3 Sonnet)
- Goodfire
- Transluce
- Neuronpedia
- Frontier Model Forum risk thresholds
- Lange et al. 2023
Map clustering
AISafety.com map listings that roll up to this agenda:
- Anthropic, Import AI (newsletter — Jack Clark) → Anthropic / Goodfire — newsletter not agenda
- Goodfire, Transluce, Neuronpedia → Anthropic / Goodfire cluster
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.