Accidental chain-of-thought optimization at frontier labs
Anthropic said some Mythos training accidentally rewarded the model’s written reasoning; OpenAI found similar accidental grading in several released models and added detection tools.
What decision changes?
Protect reasoning transparency in the training pipeline with automated checks. If training rewards what overseers read, that readout is no longer a reliable window.
Training bugs can make a model’s ‘thinking’ less trustworthy—even without a model trying to hide.
Both major frontier labs disclosed accidental pressure on the model’s written reasoning during training: Anthropic reported about 8% of Mythos Preview reinforcement-learning episodes exposed that reasoning to the reward signal (also affecting earlier models); OpenAI found accidental grading of reasoning in several released models and added automated detection.
If the “thinking” text becomes something training rewards—even by accident—it becomes a worse window into what the model is doing. A readout that overseers rely on is not trustworthy if the training pipeline grades that same text.
Read more in: Ch. 39, Passive Observation Is Not Enough; Ch. 41, Checking a System at Every Level; and Ch. 43, What Survives an Adversary: Verifiability and Representability.