AI
AIIntermediate

Claude Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5: 20× Cheaper Than Sonnet — and a 100K-Token Price Cliff

Claude Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5 compared on Anthropic’s benchmarks, list price and monthly cost for three real workloads, the 100K-token price cliff that only Haiku has, what breaks when you migrate from Haiku 4.5, and how to let Opus plan while Haiku does the work.

Cover image: Claude Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5: 20× Cheaper Than Sonnet — and a 100K-Token Price Cliff
Contents

Anthropic released Claude Haiku 5.5 on October 7, and the two numbers in the launch post that will get the most attention are a price and a jump. The price: $0.10 per million input tokens and $0.50 per million output tokens, a twentieth of Sonnet 5.5. The jump: on Terminal-Bench 4.0, an agentic coding benchmark, Haiku 4.5 scored 0.0% and Haiku 5.5 scores 39.2%.

The obvious reading is “a cheap Sonnet”. It isn’t one, and treating it like one is how teams will overpay or ship worse results. This guide compares Haiku 5.5 with Sonnet 5.5 and Opus 5.5 on Anthropic’s published numbers, works out what three real workloads cost per month on each model, and covers the one pricing rule that only Haiku has: above 100,000 tokens, the same request costs five times as much.

Claude Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5: $0.10 vs $2 vs $4 per million input tokens, Terminal-Bench 4.0 scores, and a 5× price step for prompts over 100K tokens

Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5 at a Glance

The specs below come from the Haiku 5.5 model page, the models overview and the pricing page.

Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5 at a Glance
Claude Haiku 5.5 Claude Sonnet 5.5 Claude Opus 5.5
API model ID claude-haiku-5-5 claude-sonnet-5-5 claude-opus-5-5
Input / output per million tokens $0.10 / $0.50 (prompts ≤ 100K) $2 / $10 $4 / $20
Prompts over 100K tokens $0.50 / $2.50 same as above same as above
Cache read per million tokens $0.01 ($0.05 over 100K) $0.10 $0.20
Batch API input / output $0.05 / $0.25 $1 / $5 $2 / $10
Context window / max output 1M / 128K 1M / 128K 1M / 128K
Thinking Adaptive; can be disabled Adaptive; lowest is between_tools Adaptive, always on
Default effort on the Claude API medium high medium
Comparative latency Fastest Fast Moderate
Knowledge cutoff June 2026 June 2026 June 2026
Anthropic’s own description “High-volume, latency-sensitive tasks such as classification, extraction, and routing” “The best combination of speed and intelligence” “Long-running agentic coding and knowledge work”

Three rows deserve a second look:

  • Cache reads are 10% of input on Haiku, 5% on Sonnet and Opus. On the day Haiku 5.5 launched, Anthropic also halved Sonnet 5.5’s cache read price to $0.10, which it says makes Sonnet 5.5 about 20% cheaper on most agentic work. Haiku’s cache read is still ten times cheaper than Sonnet’s. Why cache reads dominate agent bills is in the cache read rule for Fable and Mythos pricing.
  • Haiku 5.5 is the only one priced by prompt length. The other two charge the same per-token rate from 9K to 900K tokens. More on that below, because it changes how you should build long-document pipelines.
  • The default effort differs again. Haiku and Opus default to medium, Sonnet to high. Set output_config.effort explicitly in any A/B test.

Haiku 5.5 is available on the Claude API, Amazon Bedrock (anthropic.claude-haiku-5-5), Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic is also adding monthly API credits for Max and Team subscribers — $100 on Max 5x, $200 on Max 20x and up to $500 pooled for Team, usable on any model. That is enough to prototype every pattern in this article before you commit real traffic.

The Benchmarks: A Huge Jump, and a Clear Ceiling

These are the numbers from Anthropic’s launch post. Its table compares Haiku 5.5 with Haiku 4.5, OpenAI’s GPT-6 Luna and Sonnet 5.5. OSWorld here is the offline subset, so it isn’t directly comparable with the 80.1% Sonnet 5.5 scored on the full benchmark in its own launch post.

The Benchmarks: A Huge Jump, and a Clear Ceiling
Benchmark Haiku 5.5 Haiku 4.5 GPT-6 Luna Sonnet 5.5
Terminal-Bench 4.0 (agentic terminal coding) 39.2% 0.0% 16.4% 70.6%
FrontierCode 1.1 Main (mergeable changes) 46.4% — 42.4% 52.1% at xhigh
OSWorld 2.1, offline subset (computer use) 72.4% 15.7% 48.9% 83.9%
GDPval-AA v2.1 (knowledge work, Elo) 1620 735 1437 1840
AA-Briefcase v1.1 (knowledge work, Elo) 1578 614 1336 1824
Humanity’s Last Exam, no tools / with tools 45.9% / 57.4% 10.2% / 18.7% — 56.9% / 64.5%
Chartography, no tools (visual reasoning) 46.4% 6.4% 29.1% 61.6%
Benchmark comparison of Claude Haiku 4.5, Haiku 5.5 and Sonnet 5.5: Haiku 5.5 reaches 86% of Sonnet’s OSWorld score but only 56% of its Terminal-Bench 4.0 score, at 1/20 of the price
Haiku 5.5 closes most of the gap on computer use and knowledge work. On long agentic coding, it closes about half.

Three things to read out of this table:

  1. Computer use is the surprise. 72.4% on OSWorld is 86% of Sonnet’s score at 5% of its price, and Haiku 5.5 is Anthropic’s fastest model at standard speed. That combination is why Anthropic pitches it for live customer support and browser use, and why the new browser use tool supports Haiku 5.5 (Haiku 4.5 doesn’t). The toolset adds about 6,600 input tokens to every request — $0.00066 on Haiku, $0.0132 on Sonnet.
  2. Long agentic coding is where the ceiling shows. 39.2% against 70.6% on Terminal-Bench is roughly half of Sonnet’s score. Anthropic is direct about it: Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks”, while Haiku 5.5 is suited to “more narrowly scoped tasks… like compaction, summarization, or subagent work”. FrontierCode, which rewards scoped, mergeable changes, is much closer: 46.4% against 52.1%.
  3. The jump from Haiku 4.5 is the real story. GDPval-AA more than doubled and OSWorld went from 15.7% to 72.4%. The context window grows from 200K to 1M tokens, max output from 64K to 128K, and it is the first Haiku with an adjustable effort setting. If you run Haiku 4.5 today, you are upgrading — the question is only how much of your Sonnet traffic can follow it down.

The speed claim has an early customer number behind it: Asana, testing Haiku 5.5 on its AI Teammates agent, reported more than 30% lower latency for task completions and up to 2.5× faster inference per agent turn compared with the model it uses today.

As always: these are vendor-published results. Treat them as a map of where to run your own evals, not as the evals.

The 100K Cliff: Haiku’s Price Depends on Prompt Length

The pricing page states it plainly: Claude 4.6 and later models charge the same rate across the full 1M context window — except Claude Haiku 5.5, which “is priced by prompt length: a prompt of over 100,000 tokens pays higher prices.”

The 100K Cliff: Haiku’s Price Depends on Prompt Length
Haiku 5.5 rate per million tokens Prompt ≤ 100K tokens Prompt > 100K tokens
Input $0.10 $0.50
Output $0.50 $2.50
5-minute cache write $0.125 $0.625
1-hour cache write $0.20 $1.00
Cache read $0.01 $0.05
Batch input / output $0.05 / $0.25 $0.25 / $1.25

Every rate steps up by 5× — including output. A 99K-token prompt costs about $0.0099 in input; a 101K-token prompt costs about $0.0505. Anthropic says around 90% of Haiku 4.5 requests stayed under 100K, which is why it quotes “around 75% less” on average rather than the 90% the short-prompt list price suggests. The other part of that gap is the tokenizer: Haiku 5.5 uses the newer tokenizer introduced with Claude Opus 4.7, so the same text produces about 30% more tokens than on Haiku 4.5. By the docs’ rule of thumb of about 555,000 words per million tokens, 100K tokens is roughly 55,000 words of English.

The 100K-token price cliff on Claude Haiku 5.5: input cost per request rises linearly to $0.01 at 100K tokens, then jumps 5× to about $0.05 and keeps rising at the higher rate, while Sonnet 5.5 and Opus 5.5 stay flat-rate
The cliff is a step, not a slope: one token over 100K reprices the whole request.

What this means in practice:

  • Long-document pipelines should chunk below 100K. Count the whole prompt — system prompt, instructions, tool definitions and document — not just the document. The browser toolset alone is about 6,600 tokens.
  • Count tokens with the new model. A prompt you measured at 85K on Haiku 4.5 is around 110K on Haiku 5.5. Use the token counting endpoint with model: "claude-haiku-5-5".
  • Long-context work is still cheap — just not 20× cheap. Above the line, Haiku 5.5 costs $0.50 / $2.50, which is still a quarter of Sonnet 5.5’s list price.

What Three Real Workloads Cost Per Month

List prices only matter through a workload. Here are three typical ones at list prices on the Claude API. The token counts are illustrative, measured on the newer tokenizer; for Haiku 4.5, I divided them by 1.3 to account for its older tokenizer. Your mix will differ, but the ratios hold.

1. Classifying one million support tickets. A 2,000-token cached system prompt with the category definitions, a 400-token ticket, and a 30-token answer from a forced tool call — which on Haiku 5.5 returns no thinking block.

What Three Real Workloads Cost Per Month
Model Cache reads Input Output Per month
Haiku 5.5 $20 $40 $15 $75
Haiku 5.5, Batch API $10 $20 $7.50 $37.50
Haiku 4.5 (prompt too short to cache) — $1,846 $115 $1,962
Sonnet 5.5 $200 $800 $300 $1,300
Opus 5.5 $400 $1,600 $600 $2,600

Two things hide in this table. First, Haiku 4.5 can’t cache this prompt at all: its minimum cacheable prefix is 4,096 tokens, while Haiku 5.5, Sonnet 5.5 and Opus 5.5 cache from 512. Padding the Haiku 4.5 prompt past the minimum would bring it to about $833 — still 11× the Haiku 5.5 bill. Second, the table flatters Sonnet and Opus: Sonnet 5.5 rejects forced tool_choice and Opus 5.5 always thinks, so both would bill extra thinking tokens as output. Even so, Haiku 5.5 is 17× cheaper than Sonnet here. In the Batch API, cache hits are best-effort, so treat that row as a floor. For a deeper look at whether a frontier model is the right tool for classification at all, see Jev vs GPT-5 and Claude for classification and routing.

2. Summarizing 10,000 long documents. Each document is 150K tokens; the summary is 2K tokens.

What Three Real Workloads Cost Per Month
Approach Per document Per month
Haiku 5.5, one request (over 100K) $0.0800 $800
Haiku 5.5, two 75K chunks + a merge request $0.0172 $172
Haiku 4.5, one request $0.1231 $1,231
Sonnet 5.5, one request $0.32 $3,200
Opus 5.5, one request $0.64 $6,400

Chunking below the cliff makes the same job about 4.7× cheaper. The trade-off is that no single call sees the whole document — fine for summaries and extraction, wrong for questions that need cross-document reasoning.

3. A thousand agent tasks with subagents. Each task: an orchestrator reads 40K tokens and writes 4K to plan and merge, then 20 subagents each read 30K tokens of code and write 1.5K, thinking included. The orchestrator is Opus 5.5 in every row; only the subagents change.

What Three Real Workloads Cost Per Month
Subagent model Subagents per task Task total Per month
Opus 5.5 $3.00 $3.24 $3,240
Sonnet 5.5 $1.50 $1.74 $1,740
Haiku 5.5 $0.075 $0.315 $315

Ten times cheaper than all-Opus — if the subtasks are scoped well enough for a model that scores 39.2% on Terminal-Bench. That condition is the whole design problem, and it’s what the next section is about.

Opus Plans, Haiku Executes — Without Losing the Reasoning

The pattern Anthropic is selling with this release is a larger model that plans and hands subtasks to many parallel Haiku 5.5 subagents. Since Claude Fable 5.1, Anthropic restricts which models can read each other’s preserved thinking, mainly to stop distillation, so the way you wire models together matters:

  • Thinking flows up from Haiku 5.5. On the Claude API and Google Cloud, both Sonnet 5.5 and Opus 5.5 read thinking blocks from Haiku 5.5. A conversation that starts cheap on Haiku and escalates to Sonnet or Opus keeps its reasoning. On other platforms, the bigger model sees only Haiku’s text and tool calls.
  • The way back down drops reasoning. Per the docs, only Opus 5.5 reads Sonnet 5.5’s thinking, and Opus 5.5’s thinking is read only by Fable 5.1 and Mythos 5.1. A turn routed down to Haiku runs without the bigger model’s reasoning — silently, with HTTP 200, exactly like the Sonnet-Opus routing trap.
  • Thinking is bound to the account. Haiku 5.5’s thinking blocks work only in the account that produced them, or a linked one. A shared conversation store that replays sessions through another customer’s key loses them.
  • History must stay append-only. On accounts created on or after August 31, 2026, editing system, tools or earlier messages before a Haiku 5.5 thinking block returns a 400.
Opus 5.5 plans and hands scoped subtasks to Haiku 5.5 subagents in fresh conversations; escalating from Haiku to Sonnet or Opus keeps the thinking on the Claude API and Google Cloud, while routing back down to Haiku drops the bigger model’s thinking
Subagents start fresh, so nothing is dropped. Escalation keeps reasoning only in one direction.

The subagent pattern sidesteps all of this, because each subagent is a fresh conversation: the orchestrator writes a task brief, the subagent works on it, and the result comes back as text. Nothing is replayed across models, so nothing is dropped. What actually decides whether it works:

  • Brief like you’d brief a junior engineer. Inputs, the exact output format, done-criteria, and what not to touch. FrontierCode — scoped, mergeable changes — is where Haiku 5.5 comes closest to Sonnet.
  • Keep each subagent’s prompt under 100K. Twenty subagents over the line cost five times as much each.
  • Verify on the orchestrator, not in the subagent. Have the planner check results and re-dispatch, instead of letting a subagent self-certify.

If you already work through a coding agent, this is the same idea as the agent’s own subagents; how Claude Code, Codex and opencode structure that work is in AI coding agents: Claude Code vs Codex vs opencode.

Migrating From Haiku 4.5: What Breaks

The launch messaging says Haiku 5.5 is a drop-in replacement for most integrations. The migration guide is more precise — these return errors:

Migrating From Haiku 4.5: What Breaks
What you send to Haiku 4.5 today On Haiku 5.5 Fix
claude-haiku-4-5 / claude-haiku-4-5-20251001 wrong model claude-haiku-5-5 (no date suffix, no separate alias)
thinking: {"type": "enabled", "budget_tokens": N} 400 error {"type": "adaptive"} plus output_config.effort
temperature, top_p or top_k other than the defaults 400 error Omit them; steer with the prompt
A final assistant turn as prefill 400 error End with a user turn; use structured outputs or tools for format
computer_20250124 on the Claude API or Google Cloud 400 error computer_toolset_20260801
Editing earlier turns while sending thinking back 400 on accounts created on or after Aug 31, 2026 Keep the history append-only

And changes that don’t fail a request but will change your results or your bill:

  • Responses can begin with thinking blocks. Adaptive thinking is on by default. Code that reads content[0].text breaks; select blocks by type.
  • Thinking counts toward max_tokens. A small limit tuned for Haiku 4.5 can stop after the thinking block with no text at all.
  • Thinking text is empty by default. Haiku 4.5 returned summarized thinking; Haiku 5.5 returns only a signature unless you set display: "summarized".
  • The same text is about 30% more tokens. Recount prompts and recompute budgets — and recheck the 100K line.
  • No Priority Tier. If you have a Priority Tier commitment on Haiku 4.5, plan capacity separately.

A minimal classification request that works on Haiku 5.5 as written:

response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=1024,
output_config={"effort": "low"}, # for unforced calls; a forced tool call skips thinking
system=[{
"type": "text",
"text": CATEGORY_RULES, # with the ticket, keep the prompt under 100K
"cache_control": {"type": "ephemeral"}, # breakpoint on the static part, not the ticket
}],
tools=[classify_tool], # input_schema with an enum of categories
tool_choice={"type": "tool", "name": "classify"}, # forced: no thinking block
messages=[{"role": "user", "content": ticket_text}], # no assistant prefill
)
if response.stop_reason == "refusal": # no server-side fallback on Haiku 5.5
route_to_human(ticket_text)
else:
label = next(b.input for b in response.content if b.type == "tool_use")

One caching detail that silently costs money: don’t use top-level automatic caching here. It puts the breakpoint on the last block — the ticket, which changes on every request — so the cache never hits. Mark the static system block instead, and keep tools, tool_choice and effort identical across requests, because changing them invalidates the cache.

The fastest way through a real codebase is the bundled Claude API skill in Claude Code:

/claude-api migrate this project to claude-haiku-5-5

Refusals Without a Fallback

Haiku 5.5 ships with cyber and biology safeguards. Anthropic describes its cyber safeguards as more restrictive than Haiku 4.5’s but less restrictive than those on other recent models: they allow a wider range of defensive tasks than Sonnet 5.5’s, and still block penetration testing and other techniques more likely to be used by attackers. The biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5.

For engineering teams the important part is the plumbing: a declined request returns stop_reason: "refusal", and there is no server-side fallback — the work does not continue on another model. A pipeline processing a million items needs an explicit branch for that case. Teams doing authorized offensive work can apply to the Cyber Verification Program for reduced blocking.

Which One Should You Use?

  • Haiku 5.5 at low — classification, routing, extraction, tagging, moderation pre-filters, compaction and first-pass reads of documents. Force the tool call when the output is a label. Use the Batch API whenever the result isn’t needed in real time.
  • Haiku 5.5 at medium — live customer support, in-app assistants and browser use, where speed matters most and the task is well defined.
  • Haiku 5.5 as a subagent — scoped coding and research subtasks under an Opus 5.5 or Sonnet 5.5 planner, each in a fresh conversation under 100K tokens.
  • Sonnet 5.5 — your default for agentic coding and general work. It’s now 20% cheaper on agentic work thanks to the halved cache read price; the trade-offs against Opus are in Sonnet 5.5 vs Opus 5.5.
  • Opus 5.5 — ambiguous, long-horizon work and the planner seat in a multi-agent system. If you’re also weighing OpenAI’s mid-tier model, see GPT-6.1 Sol vs Claude Opus 5.5.

The Bottom Line

Haiku 5.5 is the first small Claude model that changes architecture decisions, not just bills. At 1/20 of Sonnet’s price it handles computer use, knowledge-work extraction and scoped coding well enough that a lot of traffic currently on Sonnet can move down — and almost everything still on Haiku 4.5 should move across.

But it’s a specialist, and it has a pricing rule no other current Claude model has. Keep prompts under 100,000 tokens, force tool calls for labels, give subagents fresh conversations with tight briefs, and let thinking flow up to Sonnet or Opus — it never flows back down.

Frequently asked questions

Is Claude Haiku 5.5 better than Sonnet 5.5?

No, and it isn't meant to be. On Anthropic's published table Haiku 5.5 scores 72.4% on OSWorld 2.1 (offline subset) against Sonnet 5.5's 83.9%, 46.4% on FrontierCode against 52.1%, and 39.2% on Terminal-Bench 4.0 against 70.6%. It costs 1/20 of Sonnet's list price and is Anthropic's fastest model. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding, and positions Haiku 5.5 for narrowly scoped tasks such as compaction, summarization, classification and subagent work.

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 for cache reads, $0.125 for 5-minute cache writes and $0.20 for 1-hour cache writes. For prompts over 100,000 tokens every rate is five times higher: $0.50 / $2.50, with cache reads at $0.05. The Batch API is 50% off, which brings short prompts to $0.05 / $0.25. Haiku 4.5 cost $1 / $5.

What breaks when I migrate from Claude Haiku 4.5 to Haiku 5.5?

The model ID becomes claude-haiku-5-5. Manual extended thinking with budget_tokens, non-default temperature, top_p or top_k, an assistant prefill, and the computer_20250124 tool on the Claude API and Google Cloud all return 400 errors. Responses can begin with thinking blocks, thinking text is omitted by default, thinking tokens count toward max_tokens, the same text produces about 30% more tokens, and Priority Tier is not supported. Handle stop_reason refusal, because there is no server-side fallback.

Is Claude Haiku 5.5 good for coding?

For scoped coding tasks and subagent work, yes; for long agentic coding sessions, Sonnet 5.5 or Opus 5.5 are still better. Haiku 5.5 went from 0.0% to 39.2% on Terminal-Bench 4.0 and scores 46.4% on FrontierCode 1.1, close to Sonnet 5.5's 52.1% at xhigh effort. A common pattern is an Opus 5.5 or Sonnet 5.5 orchestrator that plans the work and hands well-defined subtasks to Haiku 5.5 subagents.

How do I make Claude Haiku 5.5 as fast and cheap as possible?

Set output_config.effort to low, because adaptive thinking is on by default and thinking tokens are billed as output. You can turn thinking off with thinking type disabled at high effort or below. For classification, a forced tool_choice with an enum field works well: Haiku 5.5 accepts it, and the response starts with the tool call without a thinking block. Keep prompts under 100,000 tokens and use the Batch API for anything that isn't real time.

From the community

Discussion on Mastodon & Bluesky

Replies to this article from the open web, pulled in live. No ads, no tracking scripts.

Loading replies …
Reply on MastodonReply on Bluesky

Reply there — it shows up here automatically.

ENDE