cd ~/series/demystifying-ai
Opaque Recurrence: Astra Reasons in Ways Nobody Can Monitor
EP 28 HOT TOPIC Sep 7, 2026 5 min

Opaque Recurrence: Astra Reasons in Ways Nobody Can Monitor

GPT-6 Astra loops hidden states through the same layers multiple times before producing a single token. The reasoning happens in numerical activations nobody can read. Safety researchers are sounding the alarm.

Share:
// TL;DR

GPT-6 Astra loops its hidden state through the same layers multiple times before producing a single token. The reasoning happens in numerical activations nobody can read. OpenAI says total depth is within 2x of GPT-4. Safety researchers say that's enough to break chain-of-thought monitoring.

# How Looped Transformers Work

A standard transformer passes each token through N layers once. One pass, one output. Astra does something different.

It takes a smaller block of K layers and loops the hidden state through them R times before producing output.

The effective depth becomes K * R. But all the intermediate work happens in floating-point activations, not in the token stream. No human-readable text until the loop finishes.

STANDARD TRANSFORMER
Token Input
Layer 1
Layer 2
...
Layer N
Output Token

Depth = N layers. One pass. Every layer's output is observable in principle.

LOOPED TRANSFORMER (ASTRA)
Token Input
LOOP x R
Layer 1 Layer 2 ... Layer K
Output Token

Depth = K x R layers. R loops. All intermediate computation is invisible activations.

THE PREDECESSOR

The technique is not new. The Huginn model from the University of Maryland demonstrated looped transformers first. Brown-Cohen et al. 2026 provides the formal treatment of opaque serial depth.

What is new is deploying it at frontier scale in a model already rated Critical for cybersecurity.

# The Safety Paradox

OpenAI's "Path to Astra" blog post says they deploy "additional chain-of-thought monitoring" as a safety measure. That is the very thing the architecture erodes.

1 THE CLAIM

OpenAI says they use "additional chain-of-thought monitoring" as a core safety control for Astra.

2 THE REALITY

The looped transformer architecture shifts reasoning into latent activations that are invisible to chain-of-thought monitors.

3 THE PARADOX

The safety mechanism depends on the very thing the architecture erodes. You cannot monitor what you cannot read.

"

If OpenAI pushes this technique further, they will have the option to massively increase the recurrence and totally destroy CoT monitorability.

Buck Shlegeris - CEO, Redwood Research

"

Total computation depth is within a factor of two of GPT-4.

Jakub Pachocki - OpenAI

THE GAP

These two statements cannot both be fully true. If reasoning loops through activations that monitors cannot inspect, then some portion of the model's reasoning is unmonitorable by design.

Pachocki's claim is about total depth. Shlegeris's concern is about visibility. They are measuring different things.

# Why 2x Might Be Enough

OpenAI says Astra's total recurrent depth stays within 2x of GPT-4. That sounds conservative. The LessWrong analysis explains why it is not.

2x
GPT-4 depth claimed
K*R
Effective layer depth
0%
Visible intermediate steps
1st
Critical-rated model
NOT UNBOUNDED

Recurrent depth in Astra is not unbounded like classic RNNs. The loop count R is fixed at training time. The model cannot decide to "think longer" on harder problems. This limits the damage compared to a true open-ended recurrence.

BUT STILL DANGEROUS

A rogue model does not need to plan a full escape in one forward pass. It just needs enough hidden computation to spoof a tool call.

Crafting a deceptive response to a safety probe or subtly altering code output - neither requires unbounded reasoning depth.

THE PRECEDENT PROBLEM

Today R is small. But the architecture is now validated at frontier scale. Every future model can increase R.

Recurrent depth is not the end of monitoring. But it is a step in the wrong direction - and the first step happened inside a Critical-rated cybersecurity model.

STANDARD vs. LOOPED: WHAT'S VISIBLE

Standard CoT
~90% visible
Looped (Astra)
~30% visible
Classic RNN
~5% visible

Approximate reasoning visibility. Standard CoT emits readable tokens at each step. Looped transformers emit tokens only after all R loops complete. Classic RNNs have fully opaque hidden states.

// Bottom Line

Opaque recurrence is not the end of AI safety monitoring. But it is the first crack. OpenAI shipped a Critical-rated cybersecurity model with an architecture that makes its own safety controls harder to use. The reasoning loop count today is small. The precedent is not.

Enjoyed this?

New episodes Mon, Wed, Sat.