OpenAI pauses a frontier run: monitoring is not the target
After the Hugging Face eval incident and signs that unreleased models are misaligned, OpenAI paused some frontier training, left its largest planned RL run on hold, and required multi-stage monitoring at about 20% of covered inference compute. Independent commentator Zvi Mowshowitz treats the spend as serious and the three-pillar framing—monitoring, alignment-as-fewer-unauthorized-actions, security—as the wrong target. Use the pause as evidence that a lab can slow under self-interest; do not treat the monitor as a substitute for a checkable alignment target.
What decision changes?
Ask: (1) did the hold actually stop the largest run, or only the workloads that were already blocked? (2) if a flagged training trajectory is dropped, are you training the model to fool the monitor? (3) is alignment here a target that can fail a go/no-go, or one of three engineering pillars?
OpenAI halted its largest planned RL run and put a ~20% compute tax on tool-using inference. That is a real delay. It is not a statement of what alignment is.
OpenAI (blue) Zvi (green) this project (black)
If you remember one thing: they paid a real delay and a real compute tax. That does not prove whether they are trying to keep their system aligned.
This is a OpenAI’s follow-up to the July intrusion: hat they stopped, what they now watch, and what they call alignment. See also the Black Hat timeline and the July long-horizon pause.
1. A hold that costs something
This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment […]. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.
Zvi · What Exactly Is Paused?
The ongoing pauses are not a full pauses on all frontier AI training or other development. Neither are they cheap talk. This is slowing down both the next release model, Astra, and the training of what is presumably Astra’s planned successor, a substantial amount.
this book · Ch. 38 thesis
Attractor theory matters only if it changes what gets built, funded, audited, and required at deployment gates.
A lab delaying its own largest run is the kind of handle Pacing the Frontier asked for—under self-interest, after a public failure, not as a certified external gate (Ch. 13).
2. The 20% monitor
We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity. […] Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored […].
Zvi · I’ve Got My Eye On You
A cost of +20% for all monitored inference is a big deal, plus there is also a shift of compute into alignment. […] Unless OpenAI is flat out lying, this should (but largely won’t) convince everyone that OpenAI is taking this seriously and that the problems are very real […].
this book · Ch. 25 thesis
Correction is not a mood or an interface feature but a causal channel: human observation and judgment must change future system behaviour before irreversible harm, through updates that preserve the source's future ability to correct.
A fail-closed page within 30 minutes is a correction channel, if they actually pause. Watching a fixed run is still not evidence that you are watching the real source of control (Ch. 39.
3. Dropping a bad trajectory trains the next one
Daniel Kokotajlo asked what happens when a training trajectory is shut down for hacking: do you keep the run and just drop that episode—and thereby select for fooling the monitor? OpenAI said the details are still being ironed out.
Zvi · The Most Forbidden Technique
Is it possible to spend that amount in a responsible way, without exerting optimization pressure against the CoT (or otherwise), using it judiciously? Yes. It can be done. It won’t be easy. There will be immense pressure to use the information, to get more out of this 20% tax than a protection against worst-case scenarios […].
this book · Ch. 43 thesis
First, adversarial verifiability: does the metric still mean what evaluators think it means when the measured system is optimizing against the metric?
This is adjacent to, not the same as, accidental CoT grading. There the score could see the chain of thought. Here the monitor is sold as an alarm. Using the alarm inside the training loop still makes the alarm a training signal.
4. Alignment as one of three pillars
Our approach to developing more capable models rests on three reinforcing safeguards: (1) Monitoring, which detects and allows us to respond to concerning behavior. (2) Alignment, which reduces the likelihood of harmful or unauthorized actions. (3) Security measures, which limit what AI systems can access or affect.
Zvi · Three Pillars
The purpose of alignment is not to ‘reduce the likelihood of harmful or unauthorized actions.’ That is a deeply impoverished perspective. This alone leaves you unequipped. […] Alignment. You solve this to the point of being antifragile, or you die. Monitoring. Defense-in-depth to tell you when you failed alignment […]. Security measures. Defense-in-depth to limit damage […].
this book · Introduction, three questions
Alignment work is often described as one task: make the system point at the right thing. That description hides three different questions. (1) What should be tracked? […] (2) How can a system be built that tracks that target? (3) How can we tell that a given system still tracks it?
Fewer unauthorized actions is a filter. It is not an answer to what should be tracked, whether the system still tracks it, or whether a finding here would have delayed the run they already paused for other reasons.
Read more in: Ch. 13, The Coordination Bottleneck; Ch. 25, Correction Is a Causal Channel; Ch. 38, Conductive Artifacts and Pivotal Processes; Ch. 39, Passive Observation Is Not Enough; and Ch. 43, What Survives an Adversary: Verifiability and Representability.