METR
Do public capability evaluations track deployment-relevant risk under adversarial pressure and Goodhart Selection?
Introduction
METR measures dangerous autonomous and AI R&D capabilities at the frontier, producing eval-driven forecasting and risk reporting to inform policy and lab decisions.
Who carries it: Model Evaluation & Threat Research
What they aim to do. Measure dangerous capabilities to inform policy and lab decisions before systems reach deployment-relevant autonomy.
The hard question. Do public capability evaluations track deployment-relevant risk under adversarial pressure and Goodhart Selection?
What they produce. Autonomy evaluations, AI R&D evaluations, and frontier risk reporting including red-teaming of lab monitoring systems.
Key terms. Key terms include autonomous capabilities, AI R&D evals, eval-driven forecasting, and entity-based assessment.
Related field cruxes. Goodhart Selection; Inner Alignment
What they contribute. Empirical capability measurement at the frontier, including autonomy and AI R&D evaluations that shape lab and policy timelines.
How this project treats it. Capability evaluations do not establish correction-channel integrity or value-bundle transport across deployment; this project treats evals as an indirect signal for Inner Alignment rather than a direct guarantee.
Links
- METR
- Kwa et al. 2025 — Measuring AI Ability to Complete Long Tasks
- METR 2026 — Frontier Risk Report
- METR 2026 — Red-teaming Anthropic agent monitoring
- Barnes 2024 — General capability evaluations update
- Planned Obsolescence (Cotra)
Map clustering
AISafety.com map listings that roll up to this agenda:
- METR, Planned Obsolescence (Cotra) → METR / forecasting adjacent
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.