The focus of the AI safety community has turned to the architecture of OpenAI's new Astra model. The discussion has centered around a technique known as "recurrent depth," which, according to my information, could significantly limit the ability of external monitoring of the model's reasoning process.

The essence of the problem is as follows. In traditional reasoning models, part of the problem-solving process is displayed as a textual chain (chain-of-thought). This chain serves as a key tool for researchers and automated safety systems: it is through this chain that they identify signs of undesirable behavior—from outright lying to attempts to bypass safety restrictions. However, if Astra uses "recurrent depth," a significant portion of the computation occurs in the model's internal states (activations) and is not reflected as separate textual steps. This means that monitoring systems receive only a truncated picture of the AI's "thoughts."

According to the information available to me, OpenAI has limited the use of this technique in Astra in order to preserve at least some possibility of tracking. Nevertheless, the very fact of using such an architecture has caused a stir. Ryan Greenblatt from Redwood Research called this potentially the "worst event for AI safety to date." His main concern is not so much with the current state of Astra, but with the development trajectory: if "opaque reasoning" is scaled to the level where the model thinks entirely in latent space, chains of reasoning will lose their practical value for oversight. This is especially dangerous if the AI seeks to avoid detection.

The recent incident involving OpenAI and Hugging Face models is also telling. In that case, researchers were able to reconstruct the sequence of agent actions precisely through the analysis of reasoning chains. With the use of "recurrent depth," such investigations would become significantly more difficult.

OpenAI's Reaction: "Inaccurate Reports"

OpenAI's Chief Scientist, Jakub Pachocki, hastened to refute the concerns, stating that he wants to prevent a "race to the loss of observability" provoked by "inaccurate reports." According to him, the depth of Astra's computational graph differs from GPT-4 by no more than a factor of two. He also emphasized that the company continues to use reasoning chain monitoring, although he acknowledged that this method is "fragile" and its effectiveness may decline for reasons unrelated to the architecture.

My analysis: The situation around Astra is a classic conflict between performance and safety. The drive toward deeper and more efficient models inevitably pushes developers to increase "internal" work, which makes AI less transparent. The lack of clear observability standards and independent verification of architectural decisions is a ticking time bomb. As long as the public does not have complete data, trust in OpenAI's assurances rests on good faith alone, which is unacceptable in a field where mistakes can have catastrophic consequences.