Claude Opus 5.5 matches Fable 5.1 on most tasks at 40% lower cost. Cache reads drop 60% to $0.20/M tokens. The real story: preserved thinking ties reasoning chains to the model and conversation, making industrial-scale distillation structurally harder. Thinking can no longer be turned off. Every API call produces a signed reasoning trace that only Anthropic's infrastructure can verify.
# What Is Preserved Thinking?
Distillation is the practice of extracting a model's capabilities at scale. Run thousands of prompts, collect the reasoning traces, and use them to train a cheaper model. Anthropic's September 2026 threat intelligence report says they have detected and disrupted this at industrial scale.
Preserved thinking makes this structurally harder. Every thinking block gets a cryptographic signature. When you send it back in a multi-turn conversation, the API checks two things:
The thinking block was produced by the same model or an older one. Opus 5.5 reads Opus 5 thinking blocks. Opus 5 cannot read Opus 5.5 blocks. Forward compatibility, not backward.
The system prompt, tools, and all preceding messages are unchanged since the block was generated. Edit any prior turn? The signature breaks. The conversation must be append-only.
The result: even if you intercept every thinking block, the traces are cryptographically bound to a specific conversation state. Replay them in a different context and the API rejects them. Train on them and you get signatures, not reasoning.
# What Changes in Your Code
If you are already using extended thinking with Claude, the API surface barely changes. The key difference: thinking blocks are now opaque and signed. You pass them back but you cannot read or modify them.
import anthropic client = anthropic.Anthropic() # Turn 1: model generates preserved thinking response = client.messages.create( model="claude-opus-5-5-20261001", max_tokens=16000, thinking={"type": "enabled", "budget_tokens": 10000}, messages=[{"role": "user", "content": "Analyze this code..."}] ) # The thinking block is opaque — you see type + signature # You CANNOT read its content. You CAN pass it back. thinking_block = response.content[0] # type: "thinking" text_block = response.content[1] # type: "text" # Turn 2: pass entire content back as assistant message follow_up = client.messages.create( model="claude-opus-5-5-20261001", max_tokens=16000, thinking={"type": "enabled", "budget_tokens": 10000}, messages=[ {"role": "user", "content": "Analyze this code..."}, {"role": "assistant", "content": response.content}, {"role": "user", "content": "Now refactor the auth module"} ] )
✗ Editing or reordering previous messages → 400 error (prefix mismatch)
✗ Sending thinking block to a different model → 400 error (model mismatch)
✗ Omitting thinking block from assistant turn → model loses context, quality drops
✗ Setting thinking: disabled → no longer accepted, always-on
# How Enforcement Works
Prefix mismatch returns 400 error by default. Set prefix_mismatch_behavior: "drop_block" to silently drop failed blocks instead.
Prefix check only enforced if you opt in. But Anthropic warns: later models will enforce for all accounts. Make your integration append-only now.
Always enforced. Every account. No opt-out.
# The 40% Cost Reduction
| Model | Input $/M | Output $/M | Cache Read | Context |
|---|---|---|---|---|
| Opus 5.5 | $4.00 | $20.00 | $0.20 | 200K |
| Opus 5 | $5.00 | $25.00 | $0.50 | 200K |
| GPT-6 Sol | $2.00 | $10.00 | N/A | 1.05M |
| GPT-6 Luna | $0.10 | $0.50 | N/A | 1.05M |
The cache read price is the real story for agentic workloads. In coding agents and tool-use workflows, cache reads dominate the token bill. $0.20 per million cached tokens means the effective cost of long multi-turn coding sessions drops dramatically.
Context window stays at 200K. GPT-6 Sol offers 1.05M at lower raw prices, but Anthropic is betting that cache read pricing matters more than headline rates for the workflows that drive actual spend. For a 50-turn coding session with a 100K token system prompt, Opus 5.5 cache reads cost $2.00 total. On Opus 5 that was $5.00.
# The Safety Numbers
Nearly 2,000 adversarial scenarios.
- • Thinking can't be disabled — always-on adaptive
- • Forced tool use returns an error
- • Thinking blocks tied to model + conversation
- • Legacy
computer_20251124tool rejected - • Default effort is
medium(washigh)
# Migration Checklist: Opus 5 → 5.5
thinking: disabled from API calls. Thinking is always-on. The parameter is no longer accepted. Set budget_tokens low (1024) if you want minimal thinking. tool_choice: {"type": "tool", "name": "..."} now returns an error. Restructure your prompt to guide tool selection instead. response.content as assistant turn. Both thinking and text blocks must go back. Dropping thinking blocks means the model loses its reasoning chain. computer_20251124 to computer_20260901. The legacy tool spec is rejected by 5.5. prefix_mismatch_behavior: "drop_block" first. Silently drops invalid thinking blocks instead of 400s. Good for debugging migration issues before going strict. - ● Context window still 200K. GPT-6 Sol ships with 1.05M. For workloads that need massive context, Opus 5.5 requires chunking strategies or RAG.
- ● Preserved thinking is opaque. You cannot inspect the reasoning chain. For debugging model behavior, you see the output but not the path. This is by design but frustrating for researchers.
- ● Always-on thinking adds latency. Even with budget_tokens=1024, the model still reasons before responding. Simple classification tasks that need sub-100ms responses should stay on Haiku.
- ● No vision improvements announced. Multimodal capabilities appear unchanged from Opus 5. Image and PDF understanding are stable, not enhanced.
Preserved thinking is Anthropic's structural answer to the distillation threat. Instead of just detecting and banning scrapers, they made the reasoning chains cryptographically useless outside the original conversation. Combined with 40% lower costs, 60% cheaper cache reads, and the strongest behavioral audit scores of any Claude model, Opus 5.5 is Anthropic saying: the frontier does not have to be reckless to be accessible. The forced always-on thinking and append-only conversation model will break some integrations. That is the point — security through architecture, not policy.
10,000 Agents, 88 Hours, One Millennium Prize
OpenAI deployed 10,000 coordinating agents that claimed to solve a $1M math problem in 88 hours. Then a mathematician said they got there first.
Enjoyed this?
New episodes Mon, Wed, Sat.