cd ~/series/demystifying-ai
Recirculation - Free Lunch for Frozen Transformers
EP 29 arXiv: 2608.17981 Sep 5, 2026 5 min

Recirculation: Free Lunch for Frozen Transformers

Take any frozen transformer. Add recurrence at inference time. Get 23% better perplexity and 21% better math accuracy. No retraining. No fine-tuning. The first genuinely free architectural upgrade for deployed models.

Share:
// TL;DR

Feedforward transformers process each layer once and move on. State updates are bounded by model depth. Recirculation fixes this by leaking activations from deep layers back to shallower ones during inference. No weight changes. Applied to the Gemma 3 family: 23% reduction in perplexity, 21% improvement on GSM8k, reliable gains across downstream tasks. The trade-off: serial prefill. The payoff: a better model without touching a single weight.

# The Depth Bottleneck

Standard transformers are feedforward. Token goes in, passes through N layers, comes out. Each layer refines the hidden state once. The amount of computation per token is fixed at model depth. For hard problems, that is often not enough.

Standard Transformer
Token input
↓
Layer 1 ... Layer N
↓ one pass
Output
Computation = depth (fixed)
With Recirculation
Token input
↓
Layer 1 ... Layer N Deep activations leak back to shallow layers
↓ recurrent
Output
Computation = depth x recirculations
KEY INSIGHT

Chain-of-thought adds computation by generating more tokens. Recirculation adds computation by reprocessing existing activations. Different mechanism, complementary benefits. CoT handles complex multi-step reasoning. Recirculation handles basic state tracking that feedforward passes miss.

# How Recirculation Works

The core idea is deceptively simple. After processing a token through all layers, take the activations from deep layers and mix them back into the input of shallow layers for the next token. The model sees its own "future thoughts" when processing new input.

01
Activation Leaking Deep-to-Shallow Feedback

After the forward pass on token t, activations from deep layers are mixed into the residual stream at shallow layers for token t+1. A mixing coefficient controls how much "leaked" information blends with the standard activation. The original model weights remain completely frozen.

02
Recurrence in Two Dimensions Depth + Input Step

Looping (a related technique) recirculates in depth only - reprocessing the same token. Recirculation operates in both depth and input step. Deep activations from token t influence shallow processing of token t+1. The model becomes a dynamical system that tracks belief states across the sequence.

03
Adaptive Variant Light Hyperparameter Tuning

The base version uses fixed mixing coefficients - zero training required. The adaptive variant lightly tunes only the mixing coefficients while keeping all model weights frozen. This gets you the best results with minimal compute overhead. Think of it as calibrating the feedback loop, not retraining the model.

PSEUDOCODE
def recirculate(model, tokens, alpha=0.1):
    """Inference with recirculation. Model weights are frozen."""
    prev_deep_state = None
    
    for t, token in enumerate(tokens):
        h = model.embed(token)
        
        for layer_idx, layer in enumerate(model.layers):
            if is_shallow(layer_idx) and prev_deep_state:
                h = (1 - alpha) * h + alpha * prev_deep_state
            h = layer(h)
        
        prev_deep_state = h  # save for next token
        yield model.head(h)

# The Results

Tested on the Gemma 3 family (frozen, off-the-shelf). No retraining. No fine-tuning. Just recirculation applied at inference time.

-23%
Perplexity Reduction
+21%
GSM8k Accuracy
0
Weights Modified
Gemma 3
Model Family
Frozen
All Weights
Reliable
Downstream Gains
Standard Inference
Perplexity (suite) baseline
GSM8k baseline
Generation Latency baseline
With Recirculation
Perplexity (suite) -23%
GSM8k +21%
Generation Latency +0ms
THE TRADE-OFF

Prefill becomes serial. Standard transformers process all input tokens in parallel during prefill. Recirculation needs the deep activations from token t before processing token t+1, so prefill cannot be parallelized. But generation phase has zero additional latency - recirculation uses activations already computed. For generation-heavy workloads (chatbots, coding agents), this trade-off is favorable.

# Why This Changes the Game

The AI industry has two modes: retrain or accept what you have. Recirculation introduces a third option. Architecturally upgrade a frozen model at inference time. This is not prompt engineering. This is not fine-tuning. This is changing how the model computes without changing what it knows.

For model deployers

Take the Gemma 3 checkpoint already running in production. Add recirculation to the inference server. 23% better perplexity, zero retraining cost. Faster iteration than waiting for the next model release.

For researchers

Recirculation is complementary to looping (arXiv:2605.23872), which adds recurrence in depth only. Combining both gives recurrence in depth and across the sequence. The paper frames transformers as dynamical systems tracking belief states - a framing that opens new research directions.

For the industry

The authors describe this as "architectural evolution guided by the trained network's properties." Instead of forcing a new architecture and retraining from scratch, you observe what the trained network already does well and amplify it. Train once, upgrade the architecture later.

Recirculation
Depth + Steprecurrence
Notraining
+0msgen latency
Looping
Depth Onlyrecurrence
Notraining
+Nxgen latency
Chain-of-Thought
Token Genrecurrence
Notraining
+Nxgen latency
Fine-Tuning
Nonerecurrence
Yestraining
+0msgen latency
THE PATTERN

The trend is clear: inference-time compute is becoming the new lever. Chain-of-thought scales reasoning via tokens. Looping scales via depth passes. Recirculation scales via cross-token state tracking. All three work on frozen models. The implication: a model's effective capability at deployment time is no longer fixed at training time.

// Bottom Line

Recirculation is the rarest kind of research result: a free upgrade for models already in production. No retraining. No fine-tuning. No new weights. Just a modification to the inference loop that gives frozen transformers recurrent state tracking. 23% better perplexity. 21% better math. The trade-off (serial prefill) is real, but for generation-heavy workloads it barely matters. If this generalizes beyond Gemma 3 - and the paper's theoretical framing suggests it should - every deployed transformer just got a potential upgrade.

Enjoyed this?

New episodes Mon, Wed, Sat.