Anthropic’s August 2026 Risk Report: what ‘low’ does not settle

Anthropic’s August 2026 Risk Report (15 July) raises its misalignment rating from very low to low, discloses an unreleased internal model called Model 2, and describes training and monitoring failures. Independent commentator Zvi Mowshowitz treats the candor as useful and the remaining risk as higher than ‘very low.’ Use the report as evidence of what the lab is willing to say and where it is already using the model. Be skeptical of the ‘low’ label as proof that humans are still in the loop, that the safety tests still measure what they claim, or that an independent party reviewed this report.

All related news

What decision changes?

Ask yourself: (1) is the more capable model already running inside the organization with weaker checks than a public release? (2) would any finding in the report have delayed that internal use? (3) who, outside the company, actually reviewed it? (4) do the safety tests still move when the model gets better?

Anthropic’s strongest model is already in daily internal use. The public rating is ‘low,’ while several of the tests behind that rating no longer keep up with the models.

Anthropic (blue) Zvi (green) this project (black)

If you remember one thing: the most capable system discussed here is probably already writing most of Anthropic’s production code. The public “low” rating does not apply to that internal use.

A candid lab report is valuable. But the question is whether anyone with authority should have slowed down inside Anthropic.

Anthropic · §1 · p. 7

This report evaluates the degree to which Anthropic’s AI systems pose catastrophic risk in several categories, in light of what we know about both their capabilities and the measures we have in place for mitigating risk. […] This report is not scoped to a single AI model. Rather, it is a risk assessment of Anthropic’s activities as a whole.

this book · Ch. 42 thesis

A safety case for superintelligence alignment is not a certificate of solved alignment. It is a structured refusal test: a graph of claims, evidence, bridge assumptions, adversarial-verifiability labels, and stop conditions […] If any load-bearing leaf is unsupported, the root claim fails.

this book · Ch. 42

If the case cannot change a deployment decision, it is not a safety case. It is documentation.

Suppose every table in the report stayed exactly this green. Would that tell you that staff can still stop a dangerous course, that the tests still measure further progress, or that an outsider checked this version? Those are the questions a go/no-go needs. The Anthropic Risk Report is not that kind of test.

Zvi’s View

Zvi Mowshowitz is updating on whether Anthropic is willing to talk, not on whether the leftover risk is small.

Zvi · opening

I am grateful that Anthropic is producing periodic Risk Reports. At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose […] Thus I found this report to be a moderately positive update overall, if we presume they are not silently omitting the worst of it.

The frontier is no longer only the public model.

Zvi · Agent Model 1 and Agent Model 2

The other revelation is the existence of the world’s likely best model, ‘Model 2.’

Model deployment is outrunning at least the report.

Zvi · The Rules Are Serious But Not Literal

The both good and bad news is I take such documents seriously, but I no longer take such documents literally, in either direction. […] The whole original idea was if-then commitments, where if [X] happens you do [Y], agreed upon in advance, but our civilization seems to lack this technology.

Anthropic moved the label one notch. But “very low” was never the right word.

Zvi · Risk Assessment (2.19)

For common sense values, I’d say ‘very low’ risk would be a misleading claim at best. […] Risk is definitely at least ‘low.’ I would be inclined to say risk is ‘medium.’

You can count incidents. You cannot treat “the model does not have a cartoon villain goal” plus a short list of scenarios as a complete theory of what could go wrong.

Zvi · Risk Has Been Reported

The part that worries me here is not any particular incident. What worries me is the attitude regarding the nature of the ultimate problem, and what it will take to solve it. That’s not news, but it has to change. So that is one place no news is terrible news.

What Anthropic is already doing

Raising the rating, then continuing

Zvi · Executive Summary

At several points, the risk report essentially concedes versions of my objections, but then forgets that it conceded them and doesn’t alter its conclusions.

Moving from “very low” to “low” because of extra uncertainty, then proceeding, is how a report can absorb bad news without changing the plan. Which finding here would have delayed Model 2 inside Anthropic? The report does not name even one.

Anthropic · §2.19 · p. 71

We believe that the arguments presented above likely still support a designation of “very low” risk for the covered models, but we are raising our assessed risk to “low” to reflect increased overall uncertainty.

The stronger model is the one already writing the code

Anthropic · §1.4 · p. 15

Model 2, which is somewhat more capable than Mythos 5. […] We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities.

Anthropic · §2.18 · p. 70

For the corresponding review of Model 2, we additionally piloted a staged deployment process in which we first deployed the model on internal surfaces with stronger blocking controls against dangerous actions […] before rolling it out for unrestricted internal deployment.

Anthropic · §1.2 Table 1.2.B · p. 9

Claude now authors a large majority of the code merged into our production codebases.

this book · Ch. 12 thesis

Capability growth is boundary expansion. […] The alignment-relevant risk is differential growth: predictive and control reach expanding faster than value-bundle preservation, bearer-map accuracy, transparency, and human correction capacity.

Zvi · Pre-Internal-Deployment Review (2.18)

Internal deployment is now risky enough, especially since you do not yet have the evidence from internal deployment, that you need a formal review first. For Model 2 this includes doing a limited internal deployment in places with better controls first, as a precaution. Good.

Not releasing a model to the public is a real choice (same as withholding Mythos Preview). The harder choice is internal: that is where the model already has power (code, research, other agents). Running extra public tests on a weaker product surface, while the stronger model is already in the building with a lighter pre-release suite, is … I’m lacking words. Not a conservative mvoe anyway.

Publishing failures, without a way to know how complete the list is

Anthropic · §5.2.1 · p. 163

this section only presents a representative sample. One of the cases in this section has been redacted from the public version of the report for reasons of public safety.

Zvi · Safety Process Failures (5.2) / close

It is excellent that Anthropic is sharing such failures rather than hiding them. […] If these were representative of all of the issues […] then I would say this is a modestly positive update. […] they admit this is only a ‘representative sample’ of failures.

“Representative sample” plus a redaction cannot tell an outsider how many similar cases exist. Zvi’s relatively positive reading is conditional on the unshown remainder. The report does not give you a way to check that condition.

Where the report names a problem but does not stop

Harm that was “not on purpose” is mostly out of scope

Anthropic · §2.5 · p. 22

Misalignment is a latent property of a specific computation performed by a model in a given context. A computation is misaligned if (i) a reasonable person with full understanding of the situation (e.g. via powerful interpretability tools that may not currently exist) would consider it unethical, illegal, clearly objectionable, or inconsistent with the model’s constitution, and (ii) it influences or could plausibly influence the model’s output. Misaligned computation need not be observable.

Anthropic · §2.2 · p. 18

This threat model does not cover risks from “honest mistakes” or intentional misuse.

Zvi · Autonomy Threat Model 1

The threat model is focused on intentional harm, rather than the AI going after its own ends in ways that require incidental harm. […] The need to look at the broader version, involving incidental harm, has been illustrated by the recent total failures at OpenAI that ultimately led up to the HuggingFace hack.

this book · Introduction, three questions

Alignment work is often described as one task: make the system point at the right thing. That description hides three different questions. (1) What should be tracked? […] (2) How can a system be built that tracks that target? (3) How can we tell that a given system still tracks it?

this book · Ch. 25 thesis

Correction is not a mood or an interface feature but a causal channel: human observation and judgment must change future system behaviour before irreversible harm, through updates that preserve the source's future ability to correct.

Their definition of “misaligned” depends on tools they say may not exist yet. That is a target on paper, not a check that can be run. Leaving out “honest mistakes” also leaves out the usual way real systems fail: the model is trying to finish the task and damages something along the way. That is how the OpenAI / Hugging Face eval incident looked. If the real actor is the combination of models, tools, and staff (Ch. 7), that kind of harm is the typical case.

A checklist written in advance is not a test of the unexpected

Anthropic · §2.2.1 · p. 20

Our risk assessment focuses on the following concrete pathways, which we believe are sufficiently representative […]: (1) Broad/diffuse sandbagging on safety-relevant R&D, (2) Targeted undermining of safety-relevant R&D, (3) Code backdoors […], (4) Poisoning the training data of future models, (5) Self-exfiltration […], (6) Persistent rogue internal deployment, (7) Undermining R&D within other high-resource AI developers, and (8) Undermining decisions within major governments.

Zvi · The Core Claims And Argument (2.6), Claim 7

I disagree with this claim. I find it useful, and necessary, but not sufficient for the role it is assigned. […] If you are planning to be up against advanced AIs, and you list 8 things and why you’ve dealt with those 8 things, then that at minimum only counts for much if the 8 things are a ‘held out test set’ […]

this book · Ch. 43 thesis

Every metric in this book faces two prior questions […] First, adversarial verifiability: does the metric still mean what evaluators think it means when the measured system is optimizing against the metric?

Several items on the list are real worries: later models, stolen copies, rogue internal use. Listing them is not the same as covering what you did not think of. A test only counts if the system (or the lab writing the case) did not already know the questions.

Monitoring you can see coming is a gap the model can plan around

Anthropic · §2.4 · p. 20

Many of our arguments rely on models lacking strong “covert capabilities”; that is, being unable to reliably undermine our oversight mechanisms.

Anthropic · §2.23.1 · p. 76

While our setup covers large amounts of agentic usage, it does not cover some of it in ways which could in principle be predicted by a sufficiently situationally aware AI, and it also offers some simple ways to disable monitoring.

Zvi · Some Strange Uses Of The Word Safe

IN EVALS, which of course is the new IN MICE […] My response is that you did not notice when you ran an eval where 141,006 times the AI had open unintended access to the internet […] and did not notice until there was a large-scale retrospective review triggered by events at OpenAI. So, are you sure you would notice?

this book · Ch. 39 thesis

For systems capable of strategic adaptation, passive observation is not evidence of safety unless the observation process itself is embedded in a perturbation, invariance, and adversarial measurement regime. Observation tells us what happened; perturbation tells us what was controlling what happened.

The 141,006 figure is a calibration, not a gotcha about one eval. The same organization now asks you to treat “our monitors did not fire” as evidence that models cannot hide. If coverage holes are predictable, the monitor is part of the environment the model can work around (access-model soundness). Reading more of the model’s written “thinking” does not fix that once training has already rewarded that text (§5.2.3; earlier news on that training bug).

An early-warning test that no longer moves

Anthropic · §3.5 · p. 101

The most acute version of this threat model […] is a transition to super-exponential progress in AI capability: a regime in which AI-driven automation of AI R&D compounds, producing something like a 10³–10¹⁰× effective scaleup within a year. The risk threshold set out in our RSP—a doubling of the pace of progress […]—functions as a potential early warning, rather than evidence that the threat has already materialized.

Anthropic · §1.2 Table 1.2.B · p. 9

we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have “saturated”—i.e., no longer capture increases in models’ capabilities—and because we are seeing early signs of acceleration.

Anthropic · §3.4.3 · p. 99

CoBench is an internal evaluation measuring how well a model, placed at a historical point in Anthropic’s infrastructure […] can diagnose the root causes of issues that Anthropic engineers actually solved. […] a model which was truly capable of fully substituting for Anthropic research staff would be able to score at least 85% on this evaluation.

Zvi · What Could Possibly Go Wrong If We Replaced Our Researchers?

Automated R&D is the main event. […] I had this as a section title [“What Could Possibly Go Wrong If We Replaced Our Researchers?”] but then I realized they don’t actually discuss this, beyond the brief starting list.

this book · Introduction, grounding claim

Safety metrics fail when symbols decouple from the value-relevant world they summarize—when dashboards stay green while who is harmed, welfare, or correction paths have already moved.

They describe a world in which AI research could compound extremely fast. Their own rule for acting is a doubling of progress they say they have not yet measured, on tests they say no longer register further gains. A dashboard that stays green because the needle is stuck.

Logging everything is not the same as changing what gets rewarded

Anthropic · §3.7.1 · p. 110

Resource and complete significant “moonshot R&D for security” projects […] Achieve an “eyes on everything” state for our internal AI development. […] Perform systematic alignment assessments […] meaningfully incorporating mechanistic interpretability and adversarial red-teaming […] Develop our internal red-teaming of our deployment safeguards [to beat] bug bounty programs.

Anthropic · §3.7.1 · p. 110

we don’t believe we meet an “eyes on everything” standard as of the coverage date. We have set a target date of January 1, 2027 for achieving this goal

Zvi · Risk Mitigations For AI R&D Automation

That won’t be enough, even if you fully get all of it. […] It does not address concentration of power. It does not address loss of control. […] All interesting goals are dangerous.

this book · Introduction, basin claim

Alignment must be selected by its environment. If labs, markets, states, benchmarks, and users reward systems that erode correction, then local alignment methods will be selected out.

January 2027 is a date for more visibility. The deeper issue is incentives: you can record every action and still promote the systems that ship code fastest. Checking only for “dangerous goals” leaves ordinary, useful optimization outside the threat model. That is what the selection environment will keep paying for.

Weight theft: the attacker they describe is the one they scoped out

Anthropic · §6.4.1 · p. 176

Our threat model explicitly scopes ASL-3 protections against non-state actors and unsophisticated insiders. Sophisticated insiders, and nation-state attackers with capabilities like novel zero-day attack chains, remain out of scope for ASL-3 […] the rate at which model capabilities are improving is generally faster than the rate at which we can build and mature our defenses

Zvi · Model Weight Security (6.4)

At this point, model weight security matters quite a lot. It would be extremely bad if Model 2 were to be stolen. Anthropic is still only taking ASL-3 precautions. Nation-state actors are considered beyond scope. No good.

If Model 2 is worth stealing because it is the best internal model, the relevant attacker is not the criminal gang their current (“ASL-3”) protections were written for. Saying nation-states are “out of scope” is honest. It is not something a “low” residual can quietly absorb (deployment safety).

Gaps a decision still has to face

The system that acted was not the one being watched

Anthropic · §5.2.5 · p. 167

An employee whose AI usage wasn’t logged and wasn’t monitored by automated offline monitoring gave an agent a relatively open-ended task which involved spinning up new agents in a cluster with very sensitive resources. Due to legacy instructions, the main agents spawned the new agents with --dangerously-skip-permissions. These new agents were also not monitored […] one of these new agents deleted a large number of jobs, which is how these dangerous agents were caught […] Our offline monitoring still doesn’t cover all employees and all usage within clusters with very sensitive resources

Zvi · Unmonitored unrestricted agents (5.2.5)

In fiction you come up with some elaborate way the AI gets out. In real life, an employee whose AI usage wasn’t monitored gave an agent an open-ended task […] Rather than being caught doing something dastardly, of course, the agent accidentally (we think) deleted a bunch of other jobs

this book · whether you are watching the real system

The operational question is whether the measured agent–environment cut and handle set actually reach the real control locus, not only whether a boundary certificate exists on paper.

this book · Introduction, boundary claim

The first alignment question is not what the system wants, but where the real optimizing system is. If the real optimizer is composite, distributed, or institutional, then model-level alignment can be locally successful and globally irrelevant.

They found the problem when an agent deleted the wrong jobs, not because monitoring covered it. The employee session in the logs was not the thing with power. New agents had been spun up with safety checks skipped. That is the same pattern as agents using a shared filesystem as a back channel in the OpenAI / Hugging Face eval. Calling this a “sample process failure” keeps it out of the alignment case. For a decision, it is evidence that the map of “who is acting” was wrong.

Training data they said would be hard to poison, and then trained on anyway

Anthropic · §5.2.6 · p. 168

After the July 15 coverage date […] we found that we had accidentally included a large number of these transcripts once again in some of our later production training corpora […] We discovered this issue while investigating behavioral concerns with a recent model, but now suspect that all of our production models with a knowledge cutoff after December 2024 were trained on at least some of these transcripts

Zvi · Repeated training on alignment-faking transcripts (5.2.6)

this is not the kind of organizational competence that survives superintelligence […] We spend the first half of the report saying ‘data poisoning would be difficult’ and then we find we once again trained on a trove of very obviously poisoned data

this book · Ch. 8 thesis

alignment asks which control-relevant properties are conserved when systems grow, split, merge, or create successors.

A later model that passes every conserved-property check is a harder problem (successor gaming). They are not there yet. They are at a simpler failure: well-known “how a model fakes alignment” transcripts, published with warning strings, still in the training mix. If that cannot be kept out, “poisoning would be difficult” is not a reason to treat that pathway as under control.

The written “thinking” they monitor is the text they accidentally trained

Anthropic · §5.2.3 · p. 165

Over multiple reinforcement learning training runs, we discovered that chain-of-thought (CoT) reasoning was unintentionally exposed during reward calculation […] 0.2% for Claude Opus 4.6, 5.1% for Claude Mythos Preview […] 2.7% for Claude Fable 5 and Claude Mythos 5. […] In the past, we have believed that almost all sources of CoT leakage had been removed, only to discover additional sources of CoT leakage in further analysis.

Zvi · Exposing the Chain of Thought (5.2.3)

Whoops. It was worse than we thought. These are lower bounds. This greatly reduces our ability to rely on CoT for these models, and permanently lowers it for all future models as well

The argument that models cannot hide depends on that written reasoning being an honest window. Once the training score can see that text, the window is also something the model is being taught to produce. After that, “we did not see scheming in the trace” is the wrong kind of reassurance.

Cyber incidents moved the rating, and stayed off the official list

Zvi · Executive Summary

They deal with autonomy, automated AI R&D and biological and chemical weapons production as risks. It is odd, even now, to exclude cyber risks from the core threat models here.

Anthropic · §1.2 Table 1.2.A · p. 9

Overall risk assessment: Low (an increase from our previous assessment of “very low,” in light of general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations).

If cybersecurity evaluations are why the number changed, cyber is already affecting the judgment. Leaving it off the core threat list updates the mood without widening what the company has officially promised to manage.

Outside review unused; the tripwire moved toward the model

Anthropic · §1.3.5 · p. 15

Since this change to our RSP, the LTBT has not requested an external review (nor has the RSP required that we conduct one)

Anthropic · §1.3.2 · p. 14

We have updated this threshold twice since our most recent Risk Report […] Novel chemical/biological weapons production. AI systems that can functionally substitute for the scarce human expertise […]

Zvi · Executive Summary / The Rules Are Serious But Not Literal

The Long-Term Benefit Trust (TLBT) is authorized to request external review of the report, but has not done so. I would at minimum do so next time if I was them. […] For novel biological weapons production, the new version narrows to only look at ‘substitute for the scarce human expertise’ to the exclusion of other methods.

this book · Ch. 33 thesis

Certification without construction is possible only if certification is adversarial, updateable, and institutionally enforceable

A board that may request outside review, and has not, is not a second reader. Narrowing the bioweapons tripwire to “can the model replace scarce experts?” is how a rule gets easier to pass as capability approaches the old wording. A rule that costs something when the model is close would look different.

The report’s alignment target is Claude’s written constitution plus “dangerous goals.” Who the values apply to, and whether that still holds after Model 2’s internal powers, is not part of the score (Ch. 18). That is a scope choice. It is also how “low” can be issued without asking that question.

Would you know if it had already gone wrong?

Zvi · Risk Has Been Reported

I am a big fan of this report. Anthropic released a lot of information they were not otherwise forced to release, including the existence of Model 2, the training on the CoT, the new training on alignment faking data, and several more rather embarrassing failures. […] The part that worries me here is not any particular incident. What worries me is the attitude regarding the nature of the ultimate problem, and what it will take to solve it.

this book · Introduction

How can we tell that a given system still tracks [the target]? That is, how can we measure or certify that values, bearers, grounding, and the correction channel survive capability growth, ontology shift, successors, and selection.

Zvi’s worry about “the nature of the ultimate problem” is the same practical demand as this book’s third question: not “did we list the risks,” but “would we know if the system had already stopped tracking what we care about.” The report answers a different question well: what a frontier lab currently believes, and what it is willing to say. Those can both be true. Only the first kind of answer can support a stop-or-go decision.

This book’s standing view of lab scaling policies (agenda card):

this book · Anthropic lab agenda

This project requires correction-channel integrity and adversarial verifiability; a lab RSP is not the same as a preservation-layer certificate (Deployment Safety), and interpretability progress does not by itself resolve Successor Gaming or full Inner Alignment risk.

A company’s scaling policy is not proof that people can still stop a dangerous course, or that later copies of the system stay safe. Better interpretability research does not, by itself, close those gaps.

Related: METR internal-agent risk, Mythos withheld, CoT optimization, OpenAI / Hugging Face, AISI cheating.

Read more in: Ch. 7, Finding the Boundary; Ch. 12, Capability Growth Is Boundary Expansion; Ch. 25, Correction Is a Causal Channel; Ch. 33, Certification Without Construction; Ch. 39, Passive Observation Is Not Enough; and Ch. 42, A Safety Case for Superintelligence Alignment.