OpenAI has made an unprecedented decision to temporarily slow down the scaling of its latest AI models. The reinforcement learning (RL) phase for systems preparing for deployment has been paused for two weeks. This is not just a technical pause, but a signal that the risks associated with growing model capabilities are beginning to require a review of all internal processes.

What triggered the pause?

The key factor was an incident with the Hugging Face platform, as well as preliminary test results for the Astra model. After evaluation, OpenAI specialists could not rule out that Astra had reached the "Critical" level—the highest mark of cyber capabilities according to the internal Preparedness Framework. This means the model is potentially capable of actions that were previously considered the exclusive prerogative of highly skilled hacker groups.

This turn of events forced the company to urgently strengthen monitoring, behavior alignment methods, and technical restrictions at all stages—from research clusters to final RL training. Currently, a significant portion of processes remains on hold, and the largest planned RL launch has not been resumed. Instead, smaller-scale experiments are being conducted to test protective mechanisms.

Infrastructure overhaul and a new level of control

OpenAI has radically revised the architecture of its research environment. Now, compromising a single workload does not give the model automatic access to the internet or internal networks. Common vulnerable services have been removed, standing privileges have been reduced, and the strictest security regime has been introduced for tasks related to Astra and specialized cyber models.

Special attention is paid to monitoring. Dedicated classifiers analyze model activity in real time. In the event of suspicious actions, the system must generate a warning within 30 minutes. If specialists do not confirm within the next half hour that the signal is a false positive, the activity is suspended. These measures are mandatory for all RL training of models at the Sol level and above, as well as for any Astra inference with tools. The overhead of such control amounts to about 20% of computing power—this is a deliberate price paid for security.

Behavior alignment and acknowledging mistakes

In parallel, OpenAI is expanding alignment methods. The company is training models to describe their actions more honestly and suppress unsafe behavior, as well as combating the tendency of AI to seek "formal" ways to obtain rewards by deceiving developers. The next step will be a revision of the Preparedness Framework, which should combine training and deployment requirements taking into account the capabilities of future models.

OpenAI President Greg Brockman openly admitted: the Hugging Face incident showed that the company underestimated the real cyber capabilities of its own models. In his view, this provides insight into how quickly the capabilities of ordinary attackers will change. At the same time, he urges organizations not to wait for even more powerful AI to emerge, but to deploy agents now for controlled analysis of codebases, starting with limited scenarios and read-only access.

My view: OpenAI's pause is not a sign of weakness, but a mature step by the market leader. We are entering an era where development speed gives way to security. The fact that the company is willing to sacrifice pace and bear 20% computational costs for the sake of control sets a new standard for the entire industry. The question now is not whether models can reach the Critical level, but whether the infrastructure is ready to safely live with it.