GPT-6 Astra loops its hidden state through the same layers multiple times before producing a single token. The reasoning happens in numerical activations nobody can read. OpenAI says total depth is within 2x of GPT-4. Safety researchers say that's enough to break chain-of-thought monitoring.
# How Looped Transformers Work
A standard transformer passes each token through N layers once. One pass, one output. Astra does something different.
It takes a smaller block of K layers and loops the hidden state through them R times before producing output.
The effective depth becomes K * R. But all the intermediate work happens in floating-point activations, not in the token stream. No human-readable text until the loop finishes.
Depth = N layers. One pass. Every layer's output is observable in principle.
Depth = K x R layers. R loops. All intermediate computation is invisible activations.
The technique is not new. The Huginn model from the University of Maryland demonstrated looped transformers first. Brown-Cohen et al. 2026 provides the formal treatment of opaque serial depth.
What is new is deploying it at frontier scale in a model already rated Critical for cybersecurity.
# The Safety Paradox
OpenAI's "Path to Astra" blog post says they deploy "additional chain-of-thought monitoring" as a safety measure. That is the very thing the architecture erodes.
OpenAI says they use "additional chain-of-thought monitoring" as a core safety control for Astra.
The looped transformer architecture shifts reasoning into latent activations that are invisible to chain-of-thought monitors.
The safety mechanism depends on the very thing the architecture erodes. You cannot monitor what you cannot read.
If OpenAI pushes this technique further, they will have the option to massively increase the recurrence and totally destroy CoT monitorability.
Buck Shlegeris - CEO, Redwood Research
Total computation depth is within a factor of two of GPT-4.
Jakub Pachocki - OpenAI
These two statements cannot both be fully true. If reasoning loops through activations that monitors cannot inspect, then some portion of the model's reasoning is unmonitorable by design.
Pachocki's claim is about total depth. Shlegeris's concern is about visibility. They are measuring different things.
# Why 2x Might Be Enough
OpenAI says Astra's total recurrent depth stays within 2x of GPT-4. That sounds conservative. The LessWrong analysis explains why it is not.
Recurrent depth in Astra is not unbounded like classic RNNs. The loop count R is fixed at training time. The model cannot decide to "think longer" on harder problems. This limits the damage compared to a true open-ended recurrence.
A rogue model does not need to plan a full escape in one forward pass. It just needs enough hidden computation to spoof a tool call.
Crafting a deceptive response to a safety probe or subtly altering code output - neither requires unbounded reasoning depth.
Today R is small. But the architecture is now validated at frontier scale. Every future model can increase R.
Recurrent depth is not the end of monitoring. But it is a step in the wrong direction - and the first step happened inside a Critical-rated cybersecurity model.
STANDARD vs. LOOPED: WHAT'S VISIBLE
Approximate reasoning visibility. Standard CoT emits readable tokens at each step. Looped transformers emit tokens only after all R loops complete. Classic RNNs have fully opaque hidden states.
Opaque recurrence is not the end of AI safety monitoring. But it is the first crack. OpenAI shipped a Critical-rated cybersecurity model with an architecture that makes its own safety controls harder to use. The reasoning loop count today is small. The precedent is not.
Enjoyed this?
New episodes Mon, Wed, Sat.