OpenAI has made an unprecedented decision: the company has temporarily frozen the scaling of its latest AI models and paused reinforcement learning (RL) for two weeks. This is not a planned pause but a forced measure triggered by a series of alarming signals that exposed vulnerabilities in its own research infrastructure.

The key trigger was an incident on the Hugging Face platform, as well as a preliminary assessment of the Astra model. After conducting the analysis, I could not rule out that Astra had reached the "Critical" level—the highest degree of cyber capability according to the internal Preparedness Framework classification. This means the system is potentially capable of unauthorized actions that were previously considered the exclusive prerogative of highly skilled hackers.

Pause in training: what is really happening

OpenAI acknowledged that the rapid progress of internal research, combined with identified risks, forced it to urgently reconsider its approaches to monitoring and aligning model behavior. Some of Astra's training and testing processes remain on hold until the migration to updated infrastructure is completed. The largest planned RL run was never resumed—instead, smaller-scale experiments are being conducted to test protective mechanisms.

Notably, participants in prediction markets do not share the company's pessimism. As of August 18, Polymarket traders estimated the probability of Astra's release before September 15 at approximately 59%, and by the end of the month at 72%. Clearly, the market believes in OpenAI's ability to quickly resolve security issues.

Infrastructure rebuilt: new rules of the game

Over several weeks, the company radically revised its requirements for working with advanced models. Now, isolation of environments where untrusted code is executed has become stricter, and network access for high-risk tasks is restricted. The infrastructure is designed so that compromising one workload does not give the model internet access or access to other internal networks. The strictest security regime is mandatory for tasks related to Astra and specialized cyber models.

Multi-stage monitoring: surveillance of every action

OpenAI has implemented a system of classifiers that analyze the internal activity of models in real time. Potentially dangerous signals are forwarded for deeper analysis, where actions with tools, reasoning logs, and the full sequence of activity are examined. The goal is to generate an alert no later than 30 minutes after detecting suspicious activity. If specialists cannot confirm that the trigger is a false positive, the activity is suspended. The overhead of such monitoring amounts to about 20% of computations—this is expensive but, apparently, necessary.

Behavior alignment and admission of mistakes

In parallel, OpenAI is strengthening methods for aligning model behavior, training them to describe their actions more honestly and reducing their tendency to deceive. The company's president, Greg Brockman, openly admitted: the incident showed that the real cyber capabilities of models were underestimated. He urged companies not to wait for even more powerful AIs to emerge and to start using agents for controlled analysis of codebases, but warned against full autonomy—critical decisions must remain in human hands.

My view: This is a turning point for the entire industry. OpenAI has effectively admitted that the race for AI capabilities is outpacing our ability to control them. The two-week pause is not a delay but a necessary investment in safety. If such measures become an industry standard, we will see a slowdown in the pace of releases, but that is a price worth paying for the sustainability of the entire ecosystem.