The focus of the AI safety research community has turned to the architecture of OpenAI's new model, Astra. The discussion concerns not so much its performance as a fundamental question: to what extent are we even capable of controlling the reasoning process of such systems?
The crux of the potential problem lies in the application of a technique known as "recurrent depth." Instead of fully constructing a logical chain in the form of text, the model can perform a significant portion of its computations in internal states that are not explicitly displayed. This creates a kind of "opaque reasoning," which directly impacts the monitoring methods used to detect potentially dangerous behavior, such as bypassing safety guardrails or lying.
A Key Problem for Safety
Leading experts in AI safety have expressed concern about this. Ryan Greenblatt of Redwood Research called Astra's possible architecture one of the most serious risks to AI control. His concern is not so much about the model's current state as it is about the trajectory of its development.
The main risk, in his view, is that scaling up such an approach could lead to the creation of systems that reason entirely in latent space. This would practically deprive researchers and automated oversight systems of the ability to interpret the model's actions and intervene in a timely manner. This is especially alarming against the backdrop of recent incidents, where investigators relied precisely on textual reasoning chains to analyze the actions of AI agents.
Notably, in response to the criticism, OpenAI's Chief Scientist, Jakub Pachocki, hastened to reassure the public. He stated that the depth of Astra's computational graph does not exceed that of GPT-4 by more than a factor of two, and also emphasized that the company has always strived to maintain the ability to monitor. However, he also acknowledged that the very possibility of such control is "fragile" and may deteriorate.
This situation exposes a key paradox in AI development: the pursuit of creating ever more powerful and autonomous systems inevitably collides with the problem of their controllability. OpenAI has already assigned Astra a critical level of cyber capabilities, which implies the model's ability to independently search for vulnerabilities. The question of how much we can trust a "black box" with such capabilities becomes not merely technical, but existential.
My professional opinion: the public discussion surrounding Astra is an important signal. We stand on the threshold where model capabilities are beginning to outpace our tools for understanding them. The demand for independent evaluation and disclosure of technical details of the architecture is not idle curiosity, but a necessary measure to ensure safety at the next stage of AI development.