Anthropic / Goodfire (lab & MI stack)

Can scaling policies and interpretability keep pace with capability growth, including under strategic opacity and Goodhart Selection pressure?

Introduction

Anthropic builds frontier models under staged safety commitments (RSP), while adjacent teams and investments pursue the MI stackmechanistic interpretability research, editable internal representations, and cross-lab tooling infrastructure (Goodfire, Transluce, Neuronpedia).

Who carries it: Anthropic PBC; Goodfire; Transluce; Neuronpedia (infra); Georg Lange (causal-faithfulness seam)

What they aim to do. Build capable systems with staged safety commitments, interpretability, and editable internal representations that can be audited and adjusted.

The hard question. Can scaling policies and interpretability keep pace with capability growth, including under strategic opacity  and Goodhart Selection pressure?

What they produce. The Responsible Scaling Policy (RSP), Constitutional AI, interpretability research, Ember and editable representations, and FMF capability thresholds.

Key terms. Recurring terms include RSP, constitutional AI, mechanistic interpretability, circuits, features, steering, capability thresholds, and causal faithfulness.

Related field cruxes. Value Learning; Value Referent; Goodhart Selection; Inner Alignment; Successor Gaming; Deployment Safety

What they contribute. Industry RSP template, conditioning-predictor failure modes, Goodfire mechanistic-interpretability tooling (Anthropic investment; Apollo and DeepMind mechanistic interpretability lineage on team), and cross-lab Neuronpedia infrastructure.

How this project treats it. This project requires correction-channel integrity  and adversarial verifiability ; a lab RSP is not the same as a preservation-layer certificate (Deployment Safety), and interpretability progress does not by itself resolve Successor Gaming or full Inner Alignment risk.

Map clustering

AISafety.com map listings that roll up to this agenda:

See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.