Feedforward transformers process each layer once and move on. State updates are bounded by model depth. Recirculation fixes this by leaking activations from deep layers back to shallower ones during inference. No weight changes. Applied to the Gemma 3 family: 23% reduction in perplexity, 21% improvement on GSM8k, reliable gains across downstream tasks. The trade-off: serial prefill. The payoff: a better model without touching a single weight.
# The Depth Bottleneck
Standard transformers are feedforward. Token goes in, passes through N layers, comes out. Each layer refines the hidden state once. The amount of computation per token is fixed at model depth. For hard problems, that is often not enough.
Chain-of-thought adds computation by generating more tokens. Recirculation adds computation by reprocessing existing activations. Different mechanism, complementary benefits. CoT handles complex multi-step reasoning. Recirculation handles basic state tracking that feedforward passes miss.
# How Recirculation Works
The core idea is deceptively simple. After processing a token through all layers, take the activations from deep layers and mix them back into the input of shallow layers for the next token. The model sees its own "future thoughts" when processing new input.
After the forward pass on token t, activations from deep layers are mixed into the residual stream at shallow layers for token t+1. A mixing coefficient controls how much "leaked" information blends with the standard activation. The original model weights remain completely frozen.
Looping (a related technique) recirculates in depth only - reprocessing the same token. Recirculation operates in both depth and input step. Deep activations from token t influence shallow processing of token t+1. The model becomes a dynamical system that tracks belief states across the sequence.
The base version uses fixed mixing coefficients - zero training required. The adaptive variant lightly tunes only the mixing coefficients while keeping all model weights frozen. This gets you the best results with minimal compute overhead. Think of it as calibrating the feedback loop, not retraining the model.
def recirculate(model, tokens, alpha=0.1): """Inference with recirculation. Model weights are frozen.""" prev_deep_state = None for t, token in enumerate(tokens): h = model.embed(token) for layer_idx, layer in enumerate(model.layers): if is_shallow(layer_idx) and prev_deep_state: h = (1 - alpha) * h + alpha * prev_deep_state h = layer(h) prev_deep_state = h # save for next token yield model.head(h)
# The Results
Tested on the Gemma 3 family (frozen, off-the-shelf). No retraining. No fine-tuning. Just recirculation applied at inference time.
Prefill becomes serial. Standard transformers process all input tokens in parallel during prefill. Recirculation needs the deep activations from token t before processing token t+1, so prefill cannot be parallelized. But generation phase has zero additional latency - recirculation uses activations already computed. For generation-heavy workloads (chatbots, coding agents), this trade-off is favorable.
# Why This Changes the Game
The AI industry has two modes: retrain or accept what you have. Recirculation introduces a third option. Architecturally upgrade a frozen model at inference time. This is not prompt engineering. This is not fine-tuning. This is changing how the model computes without changing what it knows.
Take the Gemma 3 checkpoint already running in production. Add recirculation to the inference server. 23% better perplexity, zero retraining cost. Faster iteration than waiting for the next model release.
Recirculation is complementary to looping (arXiv:2605.23872), which adds recurrence in depth only. Combining both gives recurrence in depth and across the sequence. The paper frames transformers as dynamical systems tracking belief states - a framing that opens new research directions.
The authors describe this as "architectural evolution guided by the trained network's properties." Instead of forcing a new architecture and retraining from scratch, you observe what the trained network already does well and amplify it. Train once, upgrade the architecture later.
| Technique | Recurrence Type | Requires Training | Generation Latency |
|---|---|---|---|
| Recirculation | Depth + Input Step | No | +0ms |
| Looping (arXiv:2605.23872) | Depth Only | No | +Nx per loop |
| Chain-of-Thought | Token Generation | No | +Nx per token |
| Fine-Tuning | None (weight update) | Yes | +0ms |
The trend is clear: inference-time compute is becoming the new lever. Chain-of-thought scales reasoning via tokens. Looping scales via depth passes. Recirculation scales via cross-token state tracking. All three work on frozen models. The implication: a model's effective capability at deployment time is no longer fixed at training time.
Recirculation is the rarest kind of research result: a free upgrade for models already in production. No retraining. No fine-tuning. No new weights. Just a modification to the inference loop that gives frozen transformers recurrent state tracking. 23% better perplexity. 21% better math. The trade-off (serial prefill) is real, but for generation-heavy workloads it barely matters. If this generalizes beyond Gemma 3 - and the paper's theoretical framing suggests it should - every deployed transformer just got a potential upgrade.
Enjoyed this?
New episodes Mon, Wed, Sat.