Anthropic’s pace measurements: seeing the race is not winning it
Anthropic’s Institute post publishes three production-process metrics: an R&D automation index (Claude “leads” 26% as of August 2026, up from under 1% in February), oversight coverage and latency on one internal agent platform (~30,000 concurrent agents; blocked actions reviewed by humans within a week), and a one-week snapshot of safety compute share (~6% of AI R&D; ~12% of AI-driven AI R&D). They frame these as public instruments for pacing, and as possible future triggers. This book’s cut is the same as for the August risk report: a dashboard of how models are built is not a showing that human judgment still changes the next model before irreversible harm.
What decision changes?
Ask: (1) if the automation index rises, does that delay the next model’s use for further AI R&D, or only update a chart? (2) is “oversight keeping pace” coverage and flag rate, or a human judgment that changed later training, tools, or successor constraints? (3) does a week-scale human review still count when agents act in seconds? (4) when Claude scores Claude, what independent check would have counted as the index being wrong?
Claude now leads a quarter of Anthropic’s model R&D. Humans get a week to review a blocked action. Those numbers describe the race. They do not show that correction is keeping up.
Anthropic (blue) this book (black)
If you remember one thing: a public index of how much Claude builds Claude can show the race of AI vs human control. It cannot, by itself, show that humans still win it.
The August risk report already had the threat and a stuck gauge: automated R&D as the acute case, CoBench no longer resolving increments. This Institute post is the next instrument. It measures the production process instead of the task battery.
Anthropic · Reasons to track these measurements
The measurements in this piece are focused on how models are built. […] They complement capability evaluations, which measure what models can do.
Anthropic · opening
AI systems are becoming exponentially more powerful and have begun to automate more of the process of building themselves. As the world considers slowing the pace of frontier AI development, the public needs more information.
That is this project’s differential-growth claim, stated as a reporting problem.
this project · Ch. 12 thesis
The alignment-relevant risk is differential growth: predictive and control reach expanding faster than value-bundle preservation, bearer-map accuracy, transparency, and human correction capacity.
What Claude already leads
Anthropic · (1) Measuring AI-led AI R&D
As of August 2026, Claude is not operating fully autonomously for any measured subset of AI R&D work. Claude “leads” 26% of Anthropic’s AI R&D work. The share of work at or above “AI collaborates” is above 90%.
The chart they publish with that finding is the slope: under 1% in February 2026. The index is a frozen July basket of R&D tasks, judged in part by Claude.
Anthropic · Appendix · Measuring AI-led R&D
The basket is frozen. A growing index number on a July 2026 baseline tells us that the work humans were doing at that time is being automated; it does not, on its own, tell us whether new kinds of work are appearing that humans have shifted onto.
this project · Ch. 11
The central risk is not that a system becomes good at a task we already understand. The central risk is that a system becomes good at shaping the world through channels we did not classify as the task.
They name the other measurement problem too.
Anthropic · (1) Measuring AI-led AI R&D
[…] we’re using our own models to evaluate our systems, which could mean that the “judge” model could make the same kinds of errors as the model it is checking.
this project · Ch. 43 thesis
Every metric in this book faces two prior questions before it can support a safety decision. First, adversarial verifiability: does the metric still mean what evaluators think it means when the measured system is optimizing against the metric?
What the monitors cover
The agents doing that R&D are already a crowd.
Anthropic · (2) Measuring oversight of AI agents
As of August 2026, there were approximately 30,000 agents doing research and engineering work at Anthropic at any one time in our most-used internal platform. These measurements cover this platform only.
Anthropic · (2) Measuring oversight of AI agents
100% of these agents’ actions pass through the online monitor before they are executed, usually within seconds. […] Automated review happens before an agent takes action; humans review any blocked actions within one week.
That pairing is a measurement of a failure mode:
this project · Ch. 12 · Correction Lag
The system's action speed and complexity increase faster than human review. Example: an automated research agent generates and tests thousands of hypotheses per hour, but safety review remains at the pace of human papers and meetings.
Anthropic · (2) Measuring oversight of AI agents
We analyzed over a billion decisions from our research & engineering agents over August 2026 and found that 0.002% of them (about 1 in 47,000) were blocked by this monitor. […] The highest priority flags (approximately ~50 per week) are escalated to human review.
A monitor is not the correction channel this project asks for.
this project · Ch. 25 thesis
Correction is not a mood or an interface feature but a causal channel: human observation and judgment must change future system behaviour before irreversible harm, through updates that preserve the source's future ability to correct.
They offer the coverage table as the public test of whether oversight is keeping up.
Anthropic · (2) Measuring oversight of AI agents
Together, these measures would allow society to see whether oversight is keeping pace with the growing role of AI in AI R&D.
That test still needs a way to tell whether the monitor is being optimized against.
this project · Ch. 39 thesis
For systems capable of strategic adaptation, passive observation is not evidence of safety unless the observation process itself is embedded in a perturbation, invariance, and adversarial measurement regime. Observation tells us what happened; perturbation tells us what was controlling what happened.
Anthropic · Appendix · Oversight of agents
At this scale of agent activity, it is hard to be certain we have a complete picture of all possible kinds of behaviors, or that we are making these patterns adequately legible to human oversight.
Whether a number can stop the next model
The third measurement is a resource split, on one July week, with mixed work counted as R&D.
Anthropic · (3) Measuring compute allocation
Over the examined week, about 6% of compute that went to AI R&D was allocated toward safety, and about 12% of compute that went to AI-driven AI R&D was allocated toward safety.
They present all three measures as things that could later bind.
Anthropic · (1) Measuring AI-led AI R&D
These measures could also become the trigger for stronger requirements, like a fixed testing window before a new model is used for further AI R&D.
Until one of them does, this has the same issue as the risk report and the constitution: a document the lab did not have to publish, and a decision it has not shown under control.
this project · Ch. 42
If the case cannot change a deployment decision, it is not a safety case. It is documentation.
They plan to embed independent evaluators with internal-comparable access. Use the numbers as a window. Do not read them as a stop signal.
Needed work
What still has to be shown, if these measurements are to settle the race rather than describe it:
- A paired rule: what rise in the automation index, or what lag in human review, delays the next model’s use for further AI R&D.
- Uptake, not only coverage: a blocked or flagged event that changed later training, tools, or successor constraints.
- Human latency compared to irreversibility, not to a calendar week.
- A check of the index that is not Claude scoring Claude, including when new actuators sit outside the frozen basket.
Read more in: Ch. 11, Measuring Capability Without Task Ontology; Ch. 12, Capability Growth Is Boundary Expansion; Ch. 25, Correction Is a Causal Channel; Ch. 39, Passive Observation Is Not Enough; Ch. 42, A Safety Case for Superintelligence Alignment; and Ch. 43, What Survives an Adversary: Verifiability and Representability.