Two launches a week apart turned one comparison into a search trend. Anthropic released Claude Opus 5.5 on September 22, 2026 and recommends it as the starting model for most workloads. On September 29, OpenAI released GPT-6.1 Sol at $2/$10 per million tokens — exactly half of Opus 5.5’s $4/$20. The question people now type into search is short: Sol 6.1 vs Opus 5.5, which one?
The honest answer starts with a correction. These two models are not the same tier, the name “Sol” no longer means what it meant in July, and I haven’t found a benchmark that runs both on the same harness. What you can compare precisely is what each one costs you per agent turn — and that comparison has a cliff at 272,000 tokens that changes the answer.

GPT-6.1 Sol vs Claude Opus 5.5 at a Glance
The specs below come from the GPT-6.1 Sol model page, the Claude models overview and both vendors’ pricing pages.
| GPT-6.1 Sol | Claude Opus 5.5 | |
|---|---|---|
| API model ID | gpt-6.1-sol |
claude-opus-5-5 |
| Released | September 29, 2026 | September 22, 2026 |
| Vendor positioning | “Near-Astra performance for complex work at a lower cost” | “For long-running agentic coding and knowledge work” |
| Input / output per million tokens | $2 / $10 | $4 / $20 |
| Cache read per million tokens | $0.10 (5% of input) | $0.20 (5% of input) |
| Cache write per million tokens | $2.50, kept at least 30 minutes | $5 for 5 minutes · $8 for 1 hour |
| Prompts over 272K input tokens | 2× input and cache, 1.5× output for the whole request | Standard rates up to 1M |
| Context window / max output | 1.05M (922K max input) / 128K | 1M / 128K |
| Reasoning effort | low · medium (default) · high · xhigh · max |
Same five levels, medium default |
| Can reasoning be turned off? | No — none and minimal are rejected |
No — thinking is always on |
| Minimum cacheable prompt | 1,024 tokens | 512 tokens |
| Knowledge cutoff | April 30, 2026 | June 2026 |
| Tool calling | Responses API (Chat Completions without tools) | Messages API; no forced tool_choice |
| Batch discount | 50% (Flex too) | 50% |
| EU data residency on the vendor’s own API | Yes (+10% regional premium) | No — the Claude API offers us or global |
Three rows deserve a second look:
- Both vendors price cache reads at 5% of input. That sounds like parity, but 5% of $2 is half of 5% of $4. When I compared Sonnet 5.5 and Opus 5.5, identical cache-read prices shrank the real gap well below the list-price 2×. Here the opposite happens: the cheaper model is also cheaper on the line item that dominates agent bills.
- The default effort now matches. Both APIs default to
medium. The names of the five levels are the same, but they are separate scales tuned by separate vendors —highon Sol is nothighon Opus. Set effort explicitly in every request you compare. - The 272K row is the one that decides long-context work. It gets its own section below.
Which Sol? The Name Changed Tier in Three Months
If you search for “Sol vs Opus”, you will find three different models called Sol, and they don’t sit in the same place:
| Model | Released | Role in its family | Price per million tokens |
|---|---|---|---|
| GPT-5.6 Sol | July 9, 2026 | Frontier model of GPT-5.6 (above Terra and Luna) | $4 / $20 since the August 21 promotional price |
| GPT-6 Sol | September 22, 2026 | Middle of GPT-6 (below Astra) | $2 / $10, cache read $0.20 |
| GPT-6.1 Sol | September 29, 2026 | Middle of GPT-6, “near-Astra” | $2 / $10, cache read $0.10 |

Two practical consequences:
- Benchmarks labeled “Sol” may be the old flagship. Anthropic’s Opus 5.5 launch post compares against GPT-5.6 Sol — the July frontier model — not GPT-6.1 Sol. Reading those numbers as “Opus beats the new Sol” is wrong in both directions.
- GPT-6 Sol and GPT-6.1 Sol are different API models. GPT-6.1 Sol halves the cached-input price, drops the
nonereasoning effort, and requires the Responses API for tool calling. GPT-6 Sol still supportsnoneand allows function calling in Chat Completions when effort isnone. Pin the exact ID.
GPT-6.1 Sol vs Opus 5.5 isn’t a flagship fight. It’s OpenAI’s middle model against the model Anthropic tells most people to start with — and on price, Sol belongs next to Sonnet 5.5.
GPT-6.1 Sol Benchmarks: What Exists and What Doesn’t
I haven’t found GPT-6.1 Sol and Claude Opus 5.5 published side by side on the same harness, by either vendor or an independent lab. What does exist is Anthropic’s Opus 5.5 table, which includes OpenAI’s flagship GPT-6 Astra and the older GPT-5.6 Sol:
| Benchmark | Opus 5.5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic terminal coding) | 66.4% | 57.9% | 37.3% |
| FrontierCode v1.1 Main | 54.4% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | — | 41.7% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1846 | 1542 | 1588 |
| AutomationBench (business workflows) | 40.0% | 41.4% | 28.8% |
| Humanity’s Last Exam, with tools | 67.7% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 64.6% | 22.4% |
Read the table with its footnotes in mind. Opus 5.5 ran at max effort except on Terminal-Bench, where Anthropic reports its xhigh score and Astra’s high score as reported by OpenAI; those are each model’s best. Opus ran with its production safeguards on, which hand some cybersecurity and biology tasks to older models — Anthropic says this likely lowers its scores. AutomationBench was run by Zapier without fallback models, so safeguard interventions counted as failures. And every number here is vendor-published.
Anthropic also published cost-per-task claims against Astra: at its default medium effort, Opus 5.5 beats GPT-6 Astra at max effort on GDPval-AA for about a fifth of the cost per task, and matches Astra on Terminal-Bench 4.0 for about 40% of the cost.
What this tells you about Sol is an inference, not a measurement. OpenAI’s own model guide places GPT-6.1 Sol below Astra and recommends comparing the two “to assess the tradeoff between quality and cost”. On the rows where Opus 5.5 leads Astra by a wide margin — Terminal-Bench by 8.5 points, GDPval-AA by about 300 Elo, Humanity’s Last Exam by 10.5 points — a model positioned below Astra is unlikely to close the gap. On the two rows where Astra leads, Sol could lead too. Treat the table as a map of where to run your own eval, not as the eval.
GPT-6.1 Sol vs Opus 5.5 Pricing: a Full 2× Until 272K Tokens
List price says Opus 5.5 costs twice as much. Unlike the Sonnet-versus-Opus case, that holds up per agent turn, because the cache read — the line that dominates long agent loops — is also exactly half.
Take the same mid-session coding-agent turn I used for the Sonnet comparison: 60,000 tokens of cached context re-read, 3,000 new tokens written to the cache, and 1,500 output tokens including reasoning. Reasoning is billed as output on both models.
| Turn with a 60K-token context | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|
| 60K cache reads | $0.0060 | $0.0120 |
| 3K cache writes | $0.0075 | $0.0150 |
| 1.5K output incl. reasoning | $0.0150 | $0.0300 |
| Total per turn | $0.0285 | $0.0570 |
| Opus premium | — | 2.0× |
Now let the same agent work through a large codebase until each turn re-reads 400,000 tokens. On Sol, any request with more than 272K input tokens is priced at 2× input and cache rates and 1.5× output — for the whole request, not just the tokens above the threshold. Anthropic includes Opus 5.5’s full 1M-token window at standard pricing; a 900K-token request costs the same per token as a 9K one.
| Turn with a 400K-token context | GPT-6.1 Sol | Claude Opus 5.5 | Claude Sonnet 5.5 |
|---|---|---|---|
| 400K cache reads | $0.0800 | $0.0800 | $0.0800 |
| 3K cache writes | $0.0150 | $0.0150 | $0.0075 |
| 1.5K output incl. reasoning | $0.0225 | $0.0300 | $0.0150 |
| Total per turn | $0.1175 | $0.1250 | $0.1025 |
| Relative to Sol | — | 1.06× | 0.87× |

The numbers are illustrative and assume the cache-write rate doubles along with the other cache rates. Two things change your own math more than the exact token mix:
- Tokenizers differ. The same prompt is not the same number of tokens on both models. Anthropic notes that its tokenizer since Claude Opus 4.7 produces about 30% more tokens for the same text than the older one. Compare cost per task from real usage logs, never price per million tokens alone.
- Cache lifetime differs. A Sol cache write stays reusable for at least 30 minutes after its last use. An Opus 5.5 5-minute write expires after five idle minutes; to survive a human reviewing a diff, you pay the 1-hour write at $8 per million instead of $5. For agents that wait on people, that’s another line where Sol saves money.
If your agent regularly crosses 272K tokens, the price argument for Sol mostly disappears — and Sonnet 5.5 becomes the cheapest of the three, because it keeps flat pricing at Sol’s base rates.
Reasoning: Neither Model Turns It Off
Both vendors removed the “no reasoning” switch on these models, and both return an error rather than quietly ignoring the old setting:
- GPT-6.1 Sol accepts
reasoning.effortoflow,medium(default),high,xhighormax.noneandminimalare not supported; OpenAI’s migration guide says to uselowinstead. Tool calling requires the Responses API. - Claude Opus 5.5 always uses adaptive thinking.
thinking: {"type": "disabled"}or a manualbudget_tokensreturns a 400invalid_request_error;output_config.effortis the control. Anthropic also notes that Opus 5.5 thinks more per turn than Opus 5 at the same effort, most of all atxhighandmax.
A minimal request for each, with effort set explicitly so an A/B test compares like with like:
from openai import OpenAIimport anthropic
sol = OpenAI().responses.create( model="gpt-6.1-sol", reasoning={"effort": "medium"}, # "none" returns an error on this model input="Refactor the retry logic in payments/client.py",)
opus = anthropic.Anthropic().messages.create( model="claude-opus-5-5", max_tokens=16000, output_config={"effort": "medium"}, # thinking itself can't be disabled messages=[{"role": "user", "content": "Refactor the retry logic in payments/client.py"}],)Both models can also change effort mid-conversation without invalidating the cache: Sol through a configuration_update input item, Opus 5.5 through per-message effort (in beta). Use that for the “think harder on this step” case instead of switching models.
Switching Between Them Mid-Task Costs Twice
The tempting setup is a router: Sol for routine turns, Opus for the hard ones, inside the same session. Across vendors that costs you on two fronts at every switch.
- Reasoning doesn’t transfer. Each vendor’s reasoning items are its own, so the target model sees the text and tool calls but none of the reasoning behind them. Even within Anthropic’s lineup, thinking blocks are bound to the model that produced them; across vendors there’s nothing to bind.
- The cache doesn’t transfer. Prompt caches live per model and per provider. The first request on the other model re-processes the entire prefix as a cache write. For a 200K-token session, that is about $0.50 on Sol and $1.00 on Opus 5.5 before any new work — and on Sol above 272K, double that.
Route per conversation: decide at the start whether a task is Sol work or Opus work, and escalate by opening a new conversation from a summary rather than flipping models turn by turn.
For Teams in Europe
Where the model runs matters as much as what it costs if you work under GDPR or sector rules — and here the two vendors differ in a way the spec sheets don’t advertise.
- GPT-6.1 Sol supports EU data residency on OpenAI’s own API, with a 10% regional processing premium. Fast mode is not available with EU data residency.
- Claude Opus 5.5 on the Claude API offers only
usorglobalinference, and onlyusas a workspace geo. To process in Europe, use Claude on Google Cloud (regional and multi-region endpoints) or Amazon Bedrock regional endpoints, each with a 10% premium over global endpoints — and check that Opus 5.5 is offered in the region you need. Fast mode for Opus 5.5 exists only on the Claude API. - Opus 5.5 ships Anthropic’s watermarking measures for the EU AI Act and, like previous Opus models, is available with zero data retention.
If you already run on Google Cloud, Opus 5.5 through Vertex AI keeps billing, IAM and data location inside the platform you already govern. If you need first-party EU residency without a cloud intermediary, Sol is the simpler path.

Which One Should You Pick?
Pick GPT-6.1 Sol when:
- your agent loops are high-volume and cache-heavy and stay under 272K input tokens — that’s where the full 2× saving lives;
- you need EU data residency on the vendor’s own API;
- sessions pause for humans often enough that a 30-minute cache saves repeated writes;
- your team already works in Codex or the Responses API’s hosted tools.
Pick Claude Opus 5.5 when:
- the work is long-horizon and expensive to get wrong — migrations, audits, ambiguous refactors — where Anthropic’s published numbers show Opus ahead of OpenAI’s flagship;
- contexts regularly run past 272K tokens, where Sol’s discount mostly evaporates;
- your team works in Claude Code or already runs Claude on Google Cloud or Bedrock.
Consider a third option before you sign up for either: if what you like about Sol is its price, Claude Sonnet 5.5 costs the same per token, keeps flat long-context pricing and beat Opus 5.5 on Terminal-Bench 4.0. If what you like about Opus is the ceiling, GPT-6 Astra is OpenAI’s equivalent at $10/$50. And if you’re choosing the agent rather than the model, Claude Code, Codex and opencode differ in more than which model they call.
Whatever you pick, run your own eval the same way on both: identical tasks, effort set explicitly, and cost measured per completed task from usage logs.
The Bottom Line
GPT-6.1 Sol is not OpenAI’s answer to Opus 5.5; it’s OpenAI’s middle model, priced like Sonnet 5.5, with cheaper cache reads than either Claude model. For cache-heavy agent loops under 272K tokens, it costs half of Opus 5.5 per turn, and that saving is real.
Opus 5.5 is the stronger bet for the hardest long-horizon work on the evidence that exists today — Anthropic’s numbers against Astra — and its flat pricing to 1M tokens erases most of Sol’s advantage once contexts get large. Pick per workload, pin the exact model ID, set effort explicitly, and don’t switch vendors in the middle of a conversation.




From the community
Discussion on the Fediverse
Replies from Mastodon and Bluesky — straight from the open web, no tracking.
Loading replies …
No replies yet. Start the conversation:
Replies could not be loaded right now.