Anthropic withheld Claude Mythos Preview over capability risk
Anthropic announced Claude Mythos Preview with a major capability jump, declined general availability, and published an Alignment Risk Update describing process gaps and rare disallowed actions.
What decision changes?
When capability jumps, make an explicit release choice—trusted partners or limited access if needed—not automatic public deployment.
A big capability jump, especially in cyber—and Anthropic chose not to release it to the public.
Anthropic announced Claude Mythos Preview—a large capability jump, especially in cyber—and declined general public release, routing access through trusted partners (Project Glasswing). The accompanying Alignment Risk Update described rare disallowed actions, concerns that the model might notice it is being evaluated, and gaps in training and monitoring that would not be enough for more capable successors.
This is an explicit release choice: capability growth limited by misuse risk and who gets access—not only by benchmark scores.
Read more in: Ch. 12, Capability Growth Is Boundary Expansion; Ch. 33, Certification Without Construction; and Ch. 38, Conductive Artifacts and Pivotal Processes.