Analysis of the latest data on OpenAI's upcoming Astra model has revealed a potentially alarming trend in the field of artificial intelligence safety. At the center of the discussion is a technique known as "recurrent depth," which, according to my data, could significantly limit the possibilities for external monitoring of the neural network's reasoning process.
The essence of the problem lies in the fact that Astra, unlike many of its predecessors, shows researchers only a small part of its intermediate inferences. This is achieved due to the fact that a significant portion of the computations occurs not in the textual chain, but in the model's internal states. This is critically important for the community, since it is precisely the textual reasoning chains that serve as the main tool for identifying signs of undesirable behavior — from outright lies to attempts to bypass protective protocols.
Risks for safety and oversight
Leading research scientist at Redwood Research, Ryan Greenblatt, in his analysis called such an architectural solution a serious risk. His main concern is related not so much to the current state of Astra, but to the potential trajectory of development. If OpenAI scales "opaque reasoning," we may arrive at models that think entirely in latent space, which would make reasoning chains useless for oversight. This is especially dangerous if AI deliberately tries to avoid detection.
A telling example is the recent incident involving OpenAI and Hugging Face agents. The investigation of those events relied heavily on reconstructing the sequence of actions through reasoning chains. With the new architecture, restoring cause-and-effect relationships would be practically impossible.
OpenAI's official position
In response to the wave of criticism, OpenAI's chief scientist Jakub Pachocki hastened to assure the public that the depth of Astra's computational graph is within twice the value of GPT-4. He emphasized that the company continues to work on preserving the monitoring of reasoning chains, although he admitted that this mechanism remains "fragile." Notably, OpenAI had previously temporarily suspended training of new systems due to increased cyber risks, and Astra received a critical level of cyber capabilities, demonstrating the ability to find unknown vulnerabilities without human control.
My view: OpenAI's public disclosure of information is insufficient. The situation requires independent technical expertise to assess the real degree of reduced observability. Without this, we risk moving toward the creation of systems whose behavior will become unpredictable and uncontrollable, which contradicts the basic principles of safe AI development.