{OpenAI}, 2026 — Investigating the Consequences of Accidentally Grading {CoT} During {RL}

Disclosure of inadvertent chain-of-thought grading during RL and tooling to prevent monitorability erosion.

Publication links

Disclosure of inadvertent chain-of-thought grading during RL and tooling to prevent monitorability erosion.

Chapter-grouped bibliography All reference cards