The AI research community has turned its attention to OpenAI's new Astra model. The discussion is driven not by its performance, but by a potential reduction in the transparency of internal processes, which calls into question the effectiveness of existing oversight mechanisms.

The core issue: "recurrent depth"

According to information I have received, Astra employs a technique known as "recurrent depth." Unlike standard reasoning models, where logical steps are displayed as a text-based chain available for analysis, recurrent depth involves repeatedly processing the model's internal representations before generating the next textual step. This means that a significant portion of the computation occurs in the neural network's hidden states and is not explicitly reflected, potentially reducing the amount of information available to monitoring systems.

The key risk here is that researchers and automated safety systems, which rely on reasoning chains to detect signs of undesirable behavior—such as deception or attempts to bypass restrictions—may lose this critical capability. If a model can "think" in latent space without revealing its intentions in text, oversight becomes extremely difficult.

Expert assessments and OpenAI's position

This news has caused serious concern among leading specialists. Ryan Greenblatt, chief scientist at Redwood Research, called such an architectural solution potentially the "worst event for AI safety." His main concern is that scaling "opaque reasoning" could lead to models that think entirely in latent space, stripping reasoning chains of their practical value for monitoring.

In response to the criticism, OpenAI's chief scientist Jakub Pachocki hastened to assure the public that the depth of Astra's computational graph differs from GPT-4 by no more than a factor of two. He emphasized that the company aims to preserve the ability to monitor reasoning chains, although he acknowledged that this capability is "fragile" and may deteriorate for reasons unrelated to the architecture.

Despite these assurances, the situation remains concerning. It is worth recalling that in August, OpenAI already temporarily slowed the scaling of new models due to cyber risks, and Astra itself was assigned a critical level of cyber capabilities. This indicates that the company is aware of the full seriousness of potential threats.

My analysis: OpenAI's position, reducing everything to a "factor of two" relative to GPT-4, seems somewhat evasive. Even a small increase in the "depth" of computations in latent space could exponentially complicate the task of interpretation for external systems. The key question is not how much the architecture differs from the previous one, but what percentage of critically important reasoning remains beyond the visibility of monitoring. The public needs a more detailed and independent assessment of this aspect so that we can adequately evaluate the real risks.