Anthropic released Claude Haiku 5.5 on October 7, and the two numbers in the launch post that will get the most attention are a price and a jump. The price: $0.10 per million input tokens and $0.50 per million output tokens, a twentieth of Sonnet 5.5. The jump: on Terminal-Bench 4.0, an agentic coding benchmark, Haiku 4.5 scored 0.0% and Haiku 5.5 scores 39.2%.
The obvious reading is “a cheap Sonnet”. It isn’t one, and treating it like one is how teams will overpay or ship worse results. This guide compares Haiku 5.5 with Sonnet 5.5 and Opus 5.5 on Anthropic’s published numbers, works out what three real workloads cost per month on each model, and covers the one pricing rule that only Haiku has: above 100,000 tokens, the same request costs five times as much.

Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5 at a Glance
The specs below come from the Haiku 5.5 model page, the models overview and the pricing page.
| Claude Haiku 5.5 | Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|---|
| API model ID | claude-haiku-5-5 |
claude-sonnet-5-5 |
claude-opus-5-5 |
| Input / output per million tokens | $0.10 / $0.50 (prompts ≤ 100K) | $2 / $10 | $4 / $20 |
| Prompts over 100K tokens | $0.50 / $2.50 | same as above | same as above |
| Cache read per million tokens | $0.01 ($0.05 over 100K) | $0.10 | $0.20 |
| Batch API input / output | $0.05 / $0.25 | $1 / $5 | $2 / $10 |
| Context window / max output | 1M / 128K | 1M / 128K | 1M / 128K |
| Thinking | Adaptive; can be disabled | Adaptive; lowest is between_tools |
Adaptive, always on |
| Default effort on the Claude API | medium |
high |
medium |
| Comparative latency | Fastest | Fast | Moderate |
| Knowledge cutoff | June 2026 | June 2026 | June 2026 |
| Anthropic’s own description | “High-volume, latency-sensitive tasks such as classification, extraction, and routing” | “The best combination of speed and intelligence” | “Long-running agentic coding and knowledge work” |
Three rows deserve a second look:
- Cache reads are 10% of input on Haiku, 5% on Sonnet and Opus. On the day Haiku 5.5 launched, Anthropic also halved Sonnet 5.5’s cache read price to $0.10, which it says makes Sonnet 5.5 about 20% cheaper on most agentic work. Haiku’s cache read is still ten times cheaper than Sonnet’s. Why cache reads dominate agent bills is in the cache read rule for Fable and Mythos pricing.
- Haiku 5.5 is the only one priced by prompt length. The other two charge the same per-token rate from 9K to 900K tokens. More on that below, because it changes how you should build long-document pipelines.
- The default effort differs again. Haiku and Opus default to
medium, Sonnet tohigh. Setoutput_config.effortexplicitly in any A/B test.
Haiku 5.5 is available on the Claude API, Amazon Bedrock (anthropic.claude-haiku-5-5), Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic is also adding monthly API credits for Max and Team subscribers — $100 on Max 5x, $200 on Max 20x and up to $500 pooled for Team, usable on any model. That is enough to prototype every pattern in this article before you commit real traffic.
The Benchmarks: A Huge Jump, and a Clear Ceiling
These are the numbers from Anthropic’s launch post. Its table compares Haiku 5.5 with Haiku 4.5, OpenAI’s GPT-6 Luna and Sonnet 5.5. OSWorld here is the offline subset, so it isn’t directly comparable with the 80.1% Sonnet 5.5 scored on the full benchmark in its own launch post.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic terminal coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main (mergeable changes) | 46.4% | — | 42.4% | 52.1% at xhigh |
| OSWorld 2.1, offline subset (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (knowledge work, Elo) | 1578 | 614 | 1336 | 1824 |
| Humanity’s Last Exam, no tools / with tools | 45.9% / 57.4% | 10.2% / 18.7% | — | 56.9% / 64.5% |
| Chartography, no tools (visual reasoning) | 46.4% | 6.4% | 29.1% | 61.6% |

Three things to read out of this table:
- Computer use is the surprise. 72.4% on OSWorld is 86% of Sonnet’s score at 5% of its price, and Haiku 5.5 is Anthropic’s fastest model at standard speed. That combination is why Anthropic pitches it for live customer support and browser use, and why the new browser use tool supports Haiku 5.5 (Haiku 4.5 doesn’t). The toolset adds about 6,600 input tokens to every request — $0.00066 on Haiku, $0.0132 on Sonnet.
- Long agentic coding is where the ceiling shows. 39.2% against 70.6% on Terminal-Bench is roughly half of Sonnet’s score. Anthropic is direct about it: Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks”, while Haiku 5.5 is suited to “more narrowly scoped tasks… like compaction, summarization, or subagent work”. FrontierCode, which rewards scoped, mergeable changes, is much closer: 46.4% against 52.1%.
- The jump from Haiku 4.5 is the real story. GDPval-AA more than doubled and OSWorld went from 15.7% to 72.4%. The context window grows from 200K to 1M tokens, max output from 64K to 128K, and it is the first Haiku with an adjustable effort setting. If you run Haiku 4.5 today, you are upgrading — the question is only how much of your Sonnet traffic can follow it down.
The speed claim has an early customer number behind it: Asana, testing Haiku 5.5 on its AI Teammates agent, reported more than 30% lower latency for task completions and up to 2.5× faster inference per agent turn compared with the model it uses today.
As always: these are vendor-published results. Treat them as a map of where to run your own evals, not as the evals.
The 100K Cliff: Haiku’s Price Depends on Prompt Length
The pricing page states it plainly: Claude 4.6 and later models charge the same rate across the full 1M context window — except Claude Haiku 5.5, which “is priced by prompt length: a prompt of over 100,000 tokens pays higher prices.”
| Haiku 5.5 rate per million tokens | Prompt ≤ 100K tokens | Prompt > 100K tokens |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| 5-minute cache write | $0.125 | $0.625 |
| 1-hour cache write | $0.20 | $1.00 |
| Cache read | $0.01 | $0.05 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 |
Every rate steps up by 5× — including output. A 99K-token prompt costs about $0.0099 in input; a 101K-token prompt costs about $0.0505. Anthropic says around 90% of Haiku 4.5 requests stayed under 100K, which is why it quotes “around 75% less” on average rather than the 90% the short-prompt list price suggests. The other part of that gap is the tokenizer: Haiku 5.5 uses the newer tokenizer introduced with Claude Opus 4.7, so the same text produces about 30% more tokens than on Haiku 4.5. By the docs’ rule of thumb of about 555,000 words per million tokens, 100K tokens is roughly 55,000 words of English.

What this means in practice:
- Long-document pipelines should chunk below 100K. Count the whole prompt — system prompt, instructions, tool definitions and document — not just the document. The browser toolset alone is about 6,600 tokens.
- Count tokens with the new model. A prompt you measured at 85K on Haiku 4.5 is around 110K on Haiku 5.5. Use the token counting endpoint with
model: "claude-haiku-5-5". - Long-context work is still cheap — just not 20× cheap. Above the line, Haiku 5.5 costs $0.50 / $2.50, which is still a quarter of Sonnet 5.5’s list price.
What Three Real Workloads Cost Per Month
List prices only matter through a workload. Here are three typical ones at list prices on the Claude API. The token counts are illustrative, measured on the newer tokenizer; for Haiku 4.5, I divided them by 1.3 to account for its older tokenizer. Your mix will differ, but the ratios hold.
1. Classifying one million support tickets. A 2,000-token cached system prompt with the category definitions, a 400-token ticket, and a 30-token answer from a forced tool call — which on Haiku 5.5 returns no thinking block.
| Model | Cache reads | Input | Output | Per month |
|---|---|---|---|---|
| Haiku 5.5 | $20 | $40 | $15 | $75 |
| Haiku 5.5, Batch API | $10 | $20 | $7.50 | $37.50 |
| Haiku 4.5 (prompt too short to cache) | — | $1,846 | $115 | $1,962 |
| Sonnet 5.5 | $200 | $800 | $300 | $1,300 |
| Opus 5.5 | $400 | $1,600 | $600 | $2,600 |
Two things hide in this table. First, Haiku 4.5 can’t cache this prompt at all: its minimum cacheable prefix is 4,096 tokens, while Haiku 5.5, Sonnet 5.5 and Opus 5.5 cache from 512. Padding the Haiku 4.5 prompt past the minimum would bring it to about $833 — still 11× the Haiku 5.5 bill. Second, the table flatters Sonnet and Opus: Sonnet 5.5 rejects forced tool_choice and Opus 5.5 always thinks, so both would bill extra thinking tokens as output. Even so, Haiku 5.5 is 17× cheaper than Sonnet here. In the Batch API, cache hits are best-effort, so treat that row as a floor. For a deeper look at whether a frontier model is the right tool for classification at all, see Jev vs GPT-5 and Claude for classification and routing.
2. Summarizing 10,000 long documents. Each document is 150K tokens; the summary is 2K tokens.
| Approach | Per document | Per month |
|---|---|---|
| Haiku 5.5, one request (over 100K) | $0.0800 | $800 |
| Haiku 5.5, two 75K chunks + a merge request | $0.0172 | $172 |
| Haiku 4.5, one request | $0.1231 | $1,231 |
| Sonnet 5.5, one request | $0.32 | $3,200 |
| Opus 5.5, one request | $0.64 | $6,400 |
Chunking below the cliff makes the same job about 4.7× cheaper. The trade-off is that no single call sees the whole document — fine for summaries and extraction, wrong for questions that need cross-document reasoning.
3. A thousand agent tasks with subagents. Each task: an orchestrator reads 40K tokens and writes 4K to plan and merge, then 20 subagents each read 30K tokens of code and write 1.5K, thinking included. The orchestrator is Opus 5.5 in every row; only the subagents change.
| Subagent model | Subagents per task | Task total | Per month |
|---|---|---|---|
| Opus 5.5 | $3.00 | $3.24 | $3,240 |
| Sonnet 5.5 | $1.50 | $1.74 | $1,740 |
| Haiku 5.5 | $0.075 | $0.315 | $315 |
Ten times cheaper than all-Opus — if the subtasks are scoped well enough for a model that scores 39.2% on Terminal-Bench. That condition is the whole design problem, and it’s what the next section is about.
Opus Plans, Haiku Executes — Without Losing the Reasoning
The pattern Anthropic is selling with this release is a larger model that plans and hands subtasks to many parallel Haiku 5.5 subagents. Since Claude Fable 5.1, Anthropic restricts which models can read each other’s preserved thinking, mainly to stop distillation, so the way you wire models together matters:
- Thinking flows up from Haiku 5.5. On the Claude API and Google Cloud, both Sonnet 5.5 and Opus 5.5 read thinking blocks from Haiku 5.5. A conversation that starts cheap on Haiku and escalates to Sonnet or Opus keeps its reasoning. On other platforms, the bigger model sees only Haiku’s text and tool calls.
- The way back down drops reasoning. Per the docs, only Opus 5.5 reads Sonnet 5.5’s thinking, and Opus 5.5’s thinking is read only by Fable 5.1 and Mythos 5.1. A turn routed down to Haiku runs without the bigger model’s reasoning — silently, with HTTP 200, exactly like the Sonnet-Opus routing trap.
- Thinking is bound to the account. Haiku 5.5’s thinking blocks work only in the account that produced them, or a linked one. A shared conversation store that replays sessions through another customer’s key loses them.
- History must stay append-only. On accounts created on or after August 31, 2026, editing
system,toolsor earlier messages before a Haiku 5.5 thinking block returns a 400.

The subagent pattern sidesteps all of this, because each subagent is a fresh conversation: the orchestrator writes a task brief, the subagent works on it, and the result comes back as text. Nothing is replayed across models, so nothing is dropped. What actually decides whether it works:
- Brief like you’d brief a junior engineer. Inputs, the exact output format, done-criteria, and what not to touch. FrontierCode — scoped, mergeable changes — is where Haiku 5.5 comes closest to Sonnet.
- Keep each subagent’s prompt under 100K. Twenty subagents over the line cost five times as much each.
- Verify on the orchestrator, not in the subagent. Have the planner check results and re-dispatch, instead of letting a subagent self-certify.
If you already work through a coding agent, this is the same idea as the agent’s own subagents; how Claude Code, Codex and opencode structure that work is in AI coding agents: Claude Code vs Codex vs opencode.
Migrating From Haiku 4.5: What Breaks
The launch messaging says Haiku 5.5 is a drop-in replacement for most integrations. The migration guide is more precise — these return errors:
| What you send to Haiku 4.5 today | On Haiku 5.5 | Fix |
|---|---|---|
claude-haiku-4-5 / claude-haiku-4-5-20251001 |
wrong model | claude-haiku-5-5 (no date suffix, no separate alias) |
thinking: {"type": "enabled", "budget_tokens": N} |
400 error | {"type": "adaptive"} plus output_config.effort |
temperature, top_p or top_k other than the defaults |
400 error | Omit them; steer with the prompt |
| A final assistant turn as prefill | 400 error | End with a user turn; use structured outputs or tools for format |
computer_20250124 on the Claude API or Google Cloud |
400 error | computer_toolset_20260801 |
| Editing earlier turns while sending thinking back | 400 on accounts created on or after Aug 31, 2026 | Keep the history append-only |
And changes that don’t fail a request but will change your results or your bill:
- Responses can begin with thinking blocks. Adaptive thinking is on by default. Code that reads
content[0].textbreaks; select blocks bytype. - Thinking counts toward
max_tokens. A small limit tuned for Haiku 4.5 can stop after the thinking block with no text at all. - Thinking text is empty by default. Haiku 4.5 returned summarized thinking; Haiku 5.5 returns only a signature unless you set
display: "summarized". - The same text is about 30% more tokens. Recount prompts and recompute budgets — and recheck the 100K line.
- No Priority Tier. If you have a Priority Tier commitment on Haiku 4.5, plan capacity separately.
A minimal classification request that works on Haiku 5.5 as written:
response = client.messages.create( model="claude-haiku-5-5", max_tokens=1024, output_config={"effort": "low"}, # for unforced calls; a forced tool call skips thinking system=[{ "type": "text", "text": CATEGORY_RULES, # with the ticket, keep the prompt under 100K "cache_control": {"type": "ephemeral"}, # breakpoint on the static part, not the ticket }], tools=[classify_tool], # input_schema with an enum of categories tool_choice={"type": "tool", "name": "classify"}, # forced: no thinking block messages=[{"role": "user", "content": ticket_text}], # no assistant prefill)
if response.stop_reason == "refusal": # no server-side fallback on Haiku 5.5 route_to_human(ticket_text)else: label = next(b.input for b in response.content if b.type == "tool_use")One caching detail that silently costs money: don’t use top-level automatic caching here. It puts the breakpoint on the last block — the ticket, which changes on every request — so the cache never hits. Mark the static system block instead, and keep tools, tool_choice and effort identical across requests, because changing them invalidates the cache.
The fastest way through a real codebase is the bundled Claude API skill in Claude Code:
/claude-api migrate this project to claude-haiku-5-5Refusals Without a Fallback
Haiku 5.5 ships with cyber and biology safeguards. Anthropic describes its cyber safeguards as more restrictive than Haiku 4.5’s but less restrictive than those on other recent models: they allow a wider range of defensive tasks than Sonnet 5.5’s, and still block penetration testing and other techniques more likely to be used by attackers. The biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5.
For engineering teams the important part is the plumbing: a declined request returns stop_reason: "refusal", and there is no server-side fallback — the work does not continue on another model. A pipeline processing a million items needs an explicit branch for that case. Teams doing authorized offensive work can apply to the Cyber Verification Program for reduced blocking.
Which One Should You Use?
- Haiku 5.5 at
low— classification, routing, extraction, tagging, moderation pre-filters, compaction and first-pass reads of documents. Force the tool call when the output is a label. Use the Batch API whenever the result isn’t needed in real time. - Haiku 5.5 at
medium— live customer support, in-app assistants and browser use, where speed matters most and the task is well defined. - Haiku 5.5 as a subagent — scoped coding and research subtasks under an Opus 5.5 or Sonnet 5.5 planner, each in a fresh conversation under 100K tokens.
- Sonnet 5.5 — your default for agentic coding and general work. It’s now 20% cheaper on agentic work thanks to the halved cache read price; the trade-offs against Opus are in Sonnet 5.5 vs Opus 5.5.
- Opus 5.5 — ambiguous, long-horizon work and the planner seat in a multi-agent system. If you’re also weighing OpenAI’s mid-tier model, see GPT-6.1 Sol vs Claude Opus 5.5.
The Bottom Line
Haiku 5.5 is the first small Claude model that changes architecture decisions, not just bills. At 1/20 of Sonnet’s price it handles computer use, knowledge-work extraction and scoped coding well enough that a lot of traffic currently on Sonnet can move down — and almost everything still on Haiku 4.5 should move across.
But it’s a specialist, and it has a pricing rule no other current Claude model has. Keep prompts under 100,000 tokens, force tool calls for labels, give subagents fresh conversations with tight briefs, and let thinking flow up to Sonnet or Opus — it never flows back down.




From the community
Discussion on Mastodon & Bluesky
Replies to this article from the open web, pulled in live. No ads, no tracking scripts.
Reply there — it shows up here automatically.