OpenAI's Preparedness Framework flagged Astra as potentially reaching its highest 'Critical' cybersecurity capability, pausing the company's largest planned frontier AI training run for two weeks.
OpenAI paused its largest planned frontier reinforcement-learning run for two weeks after the company's own Preparedness Framework flagged an upcoming system, Astra, as potentially reaching a 'Critical' cybersecurity capability on Aug. 7. The framework, OpenAI's public system for rating model capabilities with 'Critical' as its highest tier, treats the crossing as a stop-the-run signal.
The pause covers reinforcement-learning training for deployment-bound models while OpenAI strengthens security controls and expands monitoring. Its largest planned frontier RL run remains on hold; researchers are running smaller experiments to understand model behavior and verify safeguards. The company disclosed the move on its blog, and TechCrunch reported the Astra slowdown the same day.
Monitoring is now mandatory for RL training and evaluations involving tool use at the Sol capability level and above. The pipeline starts with token-level detectors and escalates to higher-compute investigations of tool use and activity sequences, with a goal of surfacing alerts within 30 minutes. Teams must pause activity when they can't quickly rule out a flagged behavior as benign. The added observation costs about 20% of the inference compute being monitored, with the load varying by workload.
The trigger, the tier label, and the framework itself are OpenAI's own classifications under a self-administered system. The two-week clock starts now.