Anthropic released Claude Sonnet 5.5 on September 28, 2026, and one number from the launch post explains the attention it got: on Terminal-Bench 4.0, an agentic coding benchmark run in the command line, Sonnet 5.5 scores 70.6%. Opus 5.5, the flagship released a week earlier at twice the price, scores 66.4% at its best effort level.
So the obvious plan writes itself: send the routine work to Sonnet, escalate the hard parts to Opus, and pocket the difference. This guide compares the two models on the numbers Anthropic published, works out what the price gap really is per agent turn, and then shows the line in the documentation that makes that obvious router quietly worse than either model on its own.

Sonnet 5.5 vs Opus 5.5 at a Glance
The specs below come from the Sonnet 5.5 model page and the models overview.
| Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|
| API model ID | claude-sonnet-5-5 |
claude-opus-5-5 |
| Input / output per million tokens | $2 / $10 | $4 / $20 |
| Cache read per million tokens | $0.20 | $0.20 |
| 5-minute cache write | $2.50 | $5 |
| Context window / max output | 1M / 128K | 1M / 128K |
| Thinking | Adaptive; can drop to between_tools |
Adaptive, always on |
| Default effort on the Claude API | high |
medium |
| Comparative latency | Fast | Moderate |
| Knowledge cutoff | June 2026 | June 2026 |
| Anthropic’s own description | “The best combination of speed and intelligence” | “For long-running agentic coding and knowledge work” |
Two rows in that table matter more than they look:
- Cache reads cost the same. Opus 5.5 reads its cache at 5% of its input price and Sonnet 5.5 at 10% of its, which lands both at $0.20. For agent loops that re-read a long cached prefix every turn, that is most of the input bill. I covered why this matters in the cache read rule for Fable and Mythos pricing.
- The default effort differs. Sonnet 5.5 defaults to
highon the API, Opus 5.5 tomedium. If you A/B test them without settingoutput_config.effort, you are comparing Sonnet at high against Opus at medium. In Claude Code and the Claude apps, Sonnet 5.5 runs atmediumby default.
Both models are available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. A smaller Haiku 5.5 is announced “in the coming weeks”.
The Benchmarks: Where the Cheaper Model Wins
These are Anthropic’s published numbers from the launch post. The Opus 5.5 score on Terminal-Bench is its highest, measured at xhigh effort.
| Benchmark | Sonnet 5.5 | Opus 5.5 | Sonnet 5 | Sonnet vs Opus |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic terminal coding) | 70.6% | 66.4% | 10.3% | +4.2 |
| FrontierCode 1.1 Main (mergeable changes) | 52.1% at xhigh · 46.2% at max | 54.4% | 42.4% | −2.3 |
| CursorBench 4.0 (real Cursor sessions) | 55.5% | 57.8% | 34.1% | −2.3 |
| GDPval-AA v2.1 (44 occupations, Elo) | 1844 | 1846 | 1449 | −2 |
| AA-Briefcase v1.1 (knowledge work, Elo) | 1811 | 1822 | 1359 | −11 |
| Humanity’s Last Exam, with tools | 64.5% | 67.7% | 54.9% | −3.2 |
| OSWorld 2.1 (computer use) | 80.1% | 81.8% | 57.0% | −1.7 |
| Chartography, no tools | 61.6% | 64.4% | 15.6% | −2.8 |

Three things to read out of this table before you rewrite your routing config:
- The Terminal-Bench win is real but narrow. It measures multi-step work in a command line — exactly what coding agents do all day. It is not a general claim that Sonnet 5.5 is the stronger model, and Anthropic doesn’t make one: the post says Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment”, and the safety section notes that Sonnet 5.5 “doesn’t advance the frontier of our models’ capabilities”.
- More effort made one score worse. On FrontierCode, Sonnet 5.5 scored 52.1% at
xhighand only 46.2% atmax. FrontierCode penalizes changes beyond the task’s scope; at max effort Sonnet more often ran Claude Code’s code-review skill, which fans out into many subagents, and in the cases Cognition examined that led to a timeout or extra edits. Max effort is not a free upgrade. - The jump from Sonnet 5 is the bigger story. GDPval-AA rose by about 400 Elo points and OSWorld by 23 points. If you are still on Sonnet 5, you are upgrading either way; the only question is where Opus still earns its price.
Two caveats on provenance: all of these are vendor-published results, and Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment with a structured-outputs bug that Anthropic says may have slightly understated Sonnet’s scores. Treat the table as a map of where to run your own evals, not as the evals.
The Price Gap Is Not 2×
List price says Opus 5.5 costs exactly twice as much. What an agent actually pays per turn depends on its token mix, and in a long agent loop that mix is dominated by cache reads — which, as the table above shows, cost the same on both models.
Take a typical coding-agent turn in the middle of a session: 60,000 tokens of cached context re-read, 3,000 new tokens written to the cache (the last tool result and the next instruction), and 1,500 output tokens including thinking. Thinking is billed as output on both models.
| Per turn | Sonnet 5.5 | Opus 5.5 | Sonnet 5.5, twice the output |
|---|---|---|---|
| 60K cache reads × $0.20/M | $0.0120 | $0.0120 | $0.0120 |
| 3K cache writes (5-min) | $0.0075 | $0.0150 | $0.0075 |
| Output incl. thinking | $0.0150 (1.5K) | $0.0300 (1.5K) | $0.0300 (3K) |
| Total per turn | $0.0345 | $0.0570 | $0.0495 |
| Opus premium | 1.65× | — | 1.15× |

The numbers are illustrative — your token mix will differ — but the shape holds for any cache-heavy agent. The premium shrinks from 2× to about 1.65×, and if Sonnet has to think longer to reach the same answer, it shrinks further.
That is exactly what Anthropic’s own cost-per-task charts show. At low or medium effort, Sonnet 5.5 beats Sonnet 5’s best score on several benchmarks for about a tenth of the cost per task, and on FrontierCode at high effort it scores ten points above Sonnet 5 at the same setting, at about one fifteenth of the cost per task. But the launch post is candid about the top end: Sonnet 5.5 “complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost.”
In other words: the savings live at low and medium. Running Sonnet 5.5 at max to squeeze out the last points gets you roughly Opus prices without Opus judgment.
The Routing Trap: They Can’t Read Each Other’s Thinking
Here is the part that the benchmark threads miss. Since Claude Fable 5.1, Anthropic has been tightening how thinking blocks are preserved across turns, mainly to stop distillation. For Sonnet 5.5 the docs spell out which models can read whose reasoning:
- Sonnet 5.5 reads thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier models — but not from Opus 5, Opus 5.5, or any Fable or Mythos model.
- No other model reads thinking blocks from Sonnet 5.5.
When a conversation moves to a model that can’t read a block, the API drops that block before the prompt reaches the model. The request still succeeds with a 200, and the dropped blocks aren’t billed. Without an opt-in beta header, the drop is silent.

Now replay the router everyone is about to build:
- Sonnet 5.5 works through the first turns and builds up reasoning about your codebase.
- The router decides the next step is hard and sends it to Opus 5.5. Opus starts without Sonnet’s reasoning. It still sees the text and tool calls, just not the thinking behind them.
- The router drops back to Sonnet 5.5 for the easy follow-up. Sonnet doesn’t see what Opus thought.
- A cyber-related request triggers the server-side fallback to Sonnet 5. Sonnet 5 can’t read Sonnet 5.5’s reasoning either.
The same applies to any plan-on-Opus, execute-on-Sonnet setup once its Sonnet slot points at 5.5: the executor gets the plan, but not the reasoning that produced it. Dropped blocks cost nothing directly, but in the closely related prefix-check case the docs warn that Claude can sometimes think more to re-create the reasoning it lost, so a session’s token usage can still rise.
Two more bindings sit on top of the model check:
- Thinking is bound to the account. Sonnet 5.5’s thinking blocks work only in the account that produced them or a linked one. Resuming a session under a different user’s API key, shared session stores across organizations, or switching accounts mid-session in Claude Code all hit a silent drop with the reason
organization_binding_mismatch. - Thinking is bound to the conversation prefix. A thinking block stays valid only while the
systemprompt, thetoolslist and every earlier message are unchanged. For accounts created on or after August 31, 2026, 00:00 UTC, the API enforces this on Sonnet 5.5, Opus 5.5 and Fable 5.1 by default and returns a 400 on an edited history. Older accounts don’t enforce it unless asked — which means code that works on your old key can fail for a customer with a new one.
How to route without losing the plot
- Route per conversation, not per turn. Decide at the start whether a task is Sonnet work or Opus work.
- Escalate by starting fresh. When Sonnet gets stuck, summarize the task state and open a new conversation on Opus. The docs call this simple compaction and recommend it; nothing earlier is replayed, so nothing is dropped.
- Or keep Sonnet in charge and consult Opus. The advisor tool lets a Sonnet 5.5 executor consult an Opus 5.5 advisor without switching the conversation’s model. Note that the advice comes back encrypted as an
advisor_redacted_resultblock, so you can’t read it in the response. - Make drops visible. Send the
thinking-binding-controls-2026-08-01beta header and loginput_transformationson every response:
response = client.beta.messages.create( model="claude-sonnet-5-5", max_tokens=16000, thinking={ "type": "adaptive", "block_binding": {"prefix_mismatch_behavior": "drop_block"}, }, messages=messages, # append-only: never edit earlier turns betas=["thinking-binding-controls-2026-08-01"],)
for t in response.input_transformations or []: # reason: model_binding_mismatch | organization_binding_mismatch | prefix_binding_mismatch log.warning("thinking dropped at %s: %s", t.path, t.reason)On Sonnet 5.5, block_binding works only with adaptive thinking; sending it with between_tools returns a 400.
Five Breaking Changes If You Upgrade From Sonnet 5
Sonnet 5.5 has the same price as Sonnet 5 and a model ID with no date suffix, so it’s tempting to swap the string and ship. The migration guide lists five things that break:
| What you send today | On Sonnet 5.5 | Fix |
|---|---|---|
thinking: {"type": "disabled"} |
400 error | {"type": "between_tools"} — only at low, medium or high effort |
tool_choice of type any or tool |
400 error | auto plus strict: true on the tool, and say in the prompt when to call it |
computer_20251124 on the Claude API or Google Cloud |
400 error | computer_toolset_20260801 |
| Advisor tool with Sonnet 5, Opus 4.8 and some older advisors | 400 error | Pair with Opus 5, Opus 5.5, Sonnet 5.5, Fable or Mythos |
Editing earlier messages, system or tools mid-session |
400 on accounts created on or after Aug 31, 2026 | Append-only history; use mid-conversation system messages |
And three changes that don’t fail any request but will surprise you:
- Text between tool calls moves into thinking blocks. Notes longer than a sentence or two that the model writes between tool calls now come back as progress-update
thinkingblocks, which are empty at the default display setting. A UI that streams those notes to users goes quiet until you setdisplayto"updates"(beta) or"summarized", or usebetween_tools. - Effort levels are recalibrated. A level no longer produces the same amount of thinking as on Sonnet 5. Anthropic recommends starting at
mediumfor well-specified agentic coding andhighfor harder tasks — re-run your effort sweep and re-baseline cost. - Prompt caching gets cheaper to start. The minimum cacheable prompt drops from 1,024 to 512 tokens.
On Amazon Bedrock there is one extra trap: strict tool use isn’t available for Sonnet 5.5 there, so the auto plus strict replacement for forced tool use becomes auto plus your own input validation. On Google Cloud and the Claude API, strict works.
If you’re coming from Sonnet 4.6 or earlier, add the older breaking changes on top: non-default temperature, top_p or top_k return a 400, budget_tokens is gone in favour of effort, the same text produces about 30% more tokens, and a 2000×1500 image costs about 2.5 times as many tokens. From Sonnet 4.5 or earlier, prefilled assistant turns also return a 400.
The fastest way through all of it in a real codebase is the bundled Claude API skill in Claude Code:
/claude-api migrate this project to claude-sonnet-5-5A minimal request that works on Sonnet 5.5 as written:
response = client.messages.create( model="claude-sonnet-5-5", max_tokens=16000, thinking={"type": "between_tools"}, # replaces {"type": "disabled"} output_config={"effort": "medium"}, # set it explicitly; API default is high tools=[{**weather_tool, "strict": True}], tool_choice={"type": "auto"}, # "any" and "tool" now return 400 messages=messages,)
for block in response.content: # never read content[0].text blindly if block.type == "text": print(block.text)For Security Teams: The Cyber Fallback
Anthropic says Sonnet 5.5’s cyber capabilities are a large step up from Sonnet 5, so it ships with real-time safeguards similar to those on Opus 5.5 — the first Sonnet model to do so. In practice that means a decline returns stop_reason: "refusal", and stop_details can name one of five categories: cyber, bio, frontier_llm, reasoning_extraction or general_harms.
Two details matter for engineering teams:
- Higher-risk cyber work falls back to Sonnet 5. With server-side fallback enabled (
fallbacks: "default", beta, Claude API only),cyberandfrontier_llmdeclines are retried on Sonnet 5. Routine bug finding and fixing isn’t affected. Defenders who need more can apply to the Cyber Verification Program; Anthropic says expanded tiered access for Sonnet 5.5 is coming soon. And remember the routing trap: that fallback runs without Sonnet 5.5’s reasoning. - Asking for the chain of thought is now a refusal. Sonnet 5.5 is the first Sonnet to launch with classifiers against reasoning extraction. A prompt that asks the model to reproduce its internal reasoning in the answer text gets a
reasoning_extractionrefusal. Tools that “show the model’s thinking” by prompting for it need to switch todisplay: "summarized".
There is also good news for anyone running agents in containers. On Anthropic’s containment evaluations, Sonnet 5.5 came close to Opus 5.5 in how rarely it tries to escape its sandbox and was the least likely of any of their models to probe the limits of its containers. That’s a useful property, not a guarantee — recent incidents with autonomous agents are why the boundary still belongs in your infrastructure, not in the model.
Which One Should You Use?
- Sonnet 5.5 at
medium— your new default for agentic coding on well-scoped tasks, bug fixes, terminal work, CI jobs, computer use, and documents or slides. This is where the benchmark gap is smallest and the cost gap largest. - Sonnet 5.5 at
lowor withbetween_tools— chat and latency-sensitive work. For high-volume classification and extraction, a small decision model or the announced Haiku 5.5 will be cheaper still. - Opus 5.5 — ambiguous, long-horizon or open-ended work, architecture decisions, large refactors, and anything where a wrong answer costs more than the tokens. Its default
mediumeffort is already a reasonable starting point. - Sonnet 5.5 at
xhighormax— only when your own evals prove it. Cost converges on Opus, and FrontierCode shows max effort can score lower than xhigh. - Both — route per conversation, escalate through a summary or the advisor tool, and log
input_transformationsso you see every dropped block.
If you use these models through a coding agent rather than the API, the same choice shows up in /model and /effort. The agent side of that trade-off is in how AI coding agents like Claude Code, Codex and opencode actually work, and what the agent may do once it has picked a model is in Claude Code’s permission modes compared.
The Bottom Line
Sonnet 5.5 is the rare release where the cheaper model is the default answer. It matches Opus 5.5 within a couple of points on most of what Anthropic measured, beats it on terminal-based agent work, and costs half the list price — with a real per-turn gap closer to 1.65× once cache reads are counted.
But the cheapest architecture is not “Sonnet for easy turns, Opus for hard turns”. These two models don’t share their reasoning, the API won’t tell you when it drops it, and on new accounts an edited history is a hard error. Pick the model per conversation, keep that conversation append-only, and let Opus in through a fresh start or the advisor tool — not through the middle of a thought.




From the community
Discussion on the Fediverse
Replies from Mastodon and Bluesky — straight from the open web, no tracking.
Loading replies …
No replies yet. Start the conversation:
Replies could not be loaded right now.