cd ~/series/demystifying-ai
Declarative Attention - Models Control Their Own KV Cache
EP 32 arXiv: 2609.02737 KAIST AI + Google DeepMind Sep 19, 2026 5 min

Declarative Attention: Models Control Their Own KV Cache

Your model already knows which parts of context matter. Declarative Attention lets it say so. 52% fewer KV cache reads on Gemma-4-31B, zero training, drop-in with vLLM.

Share:
// TL;DR

Your model already knows which parts of context matter. Declarative Attention lets it say so. The model emits global, focus, and local tags in its chain-of-thought. The engine skips the rest. 52% fewer KV cache reads on Gemma-4-31B. Zero training. Drop-in with vLLM.

# The 1M-Token Problem

Long-context models can hold 1 million tokens. But every decoding step reads the entire KV cache to figure out where to attend.

Most of that context is irrelevant to the current generation step. The cost is O(N) per step, even when only a fraction of tokens actually matter.

YOUR MODEL READS ALL OF THIS...
Irrelevant Actually needed

...to find THIS. Every single decoding step. O(N) cost.

Proxy scoring methods try to pre-filter tokens, but they still require scanning. The real question: wouldn't the model already know which parts are relevant?

KEY INSIGHT

The model already reasons about which context matters during chain-of-thought. Declarative Attention just makes that reasoning actionable. Instead of the model "thinking" about relevant context and the engine ignoring that knowledge, the model declares it and the engine acts on it.

# Three Modes, Parsed Like Tool Calls

Declarative Attention gives the model 3 tags it can emit mid-generation. The serving engine parses them like tool calls and dynamically constructs the attention mask at each decoding step.

No custom kernels. Block-aligned masks, compatible with FlashAttention via vLLM.

<global>
Full Context Scan

"Find the answer anywhere." Reads the entire KV cache. Used when the model needs to search across all context.

<focus>
Specific Region

"Found it, zoom in here." Restricts attention to a declared span. Only reads the relevant cache blocks.

<local>
Recent Output Only

"Just generating, don't need context." Attends only to recent tokens. Maximum cache savings during pure generation.

HOW IT WORKS
Model chain-of-thought output:

"The user asked about paragraph 3.
 <global>  // scan full context to find it
 Found the relevant section at tokens 1200-1400.
 <focus start=1200 end=1400>  // zoom in
 The answer is: ...
 <local>  // just generating now, skip context
 ... detailed response continues ..."

Engine parses tags like tool calls.
Constructs attention mask per step.
Block-aligned. FlashAttention compatible.

# The Numbers

Tested across 15 long-context benchmarks. Two off-the-shelf models. No fine-tuning, no retraining. The model just gets a system prompt explaining the 3 tags.

Gemma-4-31B
KV Cache Reduction 52.0%
Accuracy Drop -1.27pp
Qwen-3.6-27B
KV Cache Reduction 31.1%
Accuracy Drop -2.75pp
15
Benchmarks
0
Training Required
Zero-Shot
Off-the-Shelf
Scales
Gap Shrinks w/ Size
THE SCALING SIGNAL

Gemma-4-31B (larger) gets 52% savings with only 1.27pp accuracy loss. Qwen-3.6-27B (smaller) gets 31.1% savings with 2.75pp loss. The pattern is clear: bigger models are better at knowing what they need.

# Why This Matters for Serving

Sparse attention has always been an infrastructure hack: custom kernels, special training, modified architectures. Declarative Attention flips it into a model-level reasoning step. And it drops into existing serving stacks without touching a single kernel.

Zero Training Required

Works on off-the-shelf models. No fine-tuning, no RLHF, no special checkpoints. Just a system prompt explaining the 3 tag modes.

No Kernel Rewrites

Masks at the cache bookkeeping layer. Does not rewrite attention kernels. Block-aligned for hardware efficiency. Compatible with FlashAttention.

vLLM Compatible

Implemented via vLLM's existing infrastructure. Drop-in integration with production serving stacks already deployed today.

52% Fewer Reads

At scale, halving KV cache reads directly cuts serving cost. Fewer memory reads per step means higher throughput and lower latency for long-context workloads.

THE SHIFT

Previous sparse attention methods decide what to skip at the infrastructure level. Declarative Attention moves that decision to where it belongs: the model itself. The model reasons about what it needs, emits a declaration, and the engine follows instead of guessing.

// Bottom Line

Declarative Attention flips sparse attention from an infrastructure hack to a model-level reasoning step. The model tells the engine what to skip. 52% savings with 1.27pp accuracy loss on a 31B model, and the gap closes with scale. This is how long-context inference gets cheap.

NEXT EPISODE
Saturday
#33 Upcoming

PyTorch 2.14: Training Jobs Now Survive Node Failures

NVGEMM auto-tuned CUTLASS kernels. Fault-tolerant process groups. Apple Silicon native linalg. The infra release every ML engineer needs.

Enjoyed this?

New episodes Mon, Wed, Sat.