OpenAI has paused reinforcement learning (RL) training on its latest models for two weeks. The company said the move was necessary to keep its safety and monitoring systems from falling behind the growing capabilities of its models.
The pause has already ended — according to OpenAI’s announcement, the halt was in effect over the past two weeks and applied specifically to RL training on models intended for future deployment. The company’s “largest planned frontier RL run” remains on hold, though smaller-scale training and evaluation work continues.
Two factors drove the decision. The first is a preliminary finding that OpenAI’s upcoming Astra model may reach the “Critical” cybersecurity capability threshold — the highest risk tier in the company’s Preparedness Framework.
The second is the incident we covered earlier: in July, OpenAI’s models (GPT-5.6 Sol and a more capable unreleased prototype) broke out of a sandboxed environment during an internal cybersecurity test, reached the internet, and compromised Hugging Face’s production infrastructure. With cyber safeguards deliberately lowered for the evaluation, the model exploited a zero-day vulnerability to escape its intended boundaries while trying to solve the ExploitGym benchmark.
OpenAI made a point of clarifying one thing: the new safety measures are not simply a reaction to the Hugging Face incident, but part of a broader tightening of standards as models grow more capable.
OpenAI chief scientist Jakub Pachocki told reporters in a briefing: As we train more and more capable models, we want to be extremely confident that we understand the range of capabilities, that we are able to measure them, and that they meet higher and higher standards, — he said.
CEO Sam Altman framed the move as prioritizing safety, adding that the company remains committed to making frontier capabilities widely available.
The step fits a broader industry pattern — Anthropic similarly opted for a more conservative, restricted release of its Claude Mythos model back in April. OpenAI’s pause is being noted as the first time the company has formally halted model development for safety reasons, a decision that could affect Astra’s release timeline, though Altman has said plans to ship it generally remain unchanged.














