Anthropic released Claude Sonnet 5.5 on September 28, 2026, six days after Opus 5.5. The API model ID is claude-sonnet-5-5. It keeps Sonnet 5’s price of $2 per million input tokens and $10 per million output, and Anthropic says it generates output 30%+ faster and costs up to 30% less per task because it needs fewer tokens.
We didn’t want to repeat the launch post, so we ran it. Same coding task, same agent, three runs each for Sonnet 5.5, Sonnet 5, Opus 5.5, and OpenAI’s GPT-6 Sol, the model Anthropic itself compares against.
Sonnet 5.5 vs Opus 5.5: which should you use?
Use Sonnet 5.5 when the task has a clear spec: bug fixes, well-defined features, documents, repeatable agent jobs. Use Opus 5.5 for long, open-ended work that needs judgment across many steps. Opus 5.5 costs exactly twice as much per token ($4/$20) and still leads on repository-scale coding, most clearly on SWE-Bench Pro (89.9% vs 81.3%).
Anthropic’s own framing matches that split. It calls Sonnet 5.5 “a faster, lower-cost complement to Claude Opus 5.5” and says it is “strongest at well-scoped everyday tasks, fixing bugs,” documents, slides, and spreadsheets. Its prompting guide is blunter: “for the hardest long-horizon work, an Opus model is the better choice.”
Here is the gap, benchmark by benchmark:
| Benchmark | Sonnet 5.5 | Opus 5.5 | Gap |
|---|---|---|---|
| SWE-Bench Pro | 81.3% | 89.9% | Opus +8.6 |
| FrontierCode v1.1 | 46.2% | 54.4% | Opus +8.2 |
| Humanity’s Last Exam (with tools) | 64.5% | 67.7% | Opus +3.2 |
| CursorBench 4.0 | 55.5% | 57.8% | Opus +2.3 |
| OSWorld 2.1 | 80.1% | 81.8% | Opus +1.7 |
| GDPval-AA v2.1 (Elo) | 1844 | 1846 | tie |
| AutomationBench | 44.7% | 42.5% | Sonnet +2.2 |
| Terminal-Bench 4.0 | 70.6% | 66.4% | Sonnet +4.2 |
Anthropic’s announcement figures. Sonnet 5.5 is shown at max effort; on FrontierCode it scores higher at xhigh (52.1%), which narrows that gap to about 2 points. Opus 5.5’s Terminal-Bench score is at xhigh (64.8% at max).
Opus keeps a clear lead on SWE-Bench Pro and a smaller one on hard reasoning. On computer use, office work, and terminal tasks, Sonnet 5.5 sits within two points or ahead, which is the ground most agents actually cover.
In our run the two models behaved almost identically: about 6 steps each, every hidden test passed, but Opus 5.5 took 68 seconds and $0.33 per run against Sonnet 5.5’s 48 seconds and $0.19. On a task this size, Opus bought nothing extra.
Sonnet 5.5 vs Sonnet 5: is the upgrade worth it?
Yes. The per-token price is identical, Sonnet 5.5 scores higher on every benchmark in Anthropic’s table, and in our test it did the same job 2.6x faster for 37% less money. The only reasons to wait are the new cyber safeguards and a handful of API changes that break old request code, both covered below.
The jumps are large for a point release:
| Benchmark | Sonnet 5.5 | Sonnet 5 |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% |
| OSWorld 2.1 (partial score) | 80.1% | 57.0% |
| SWE-Bench Pro | 81.3% | 63.2% |
| CursorBench 4.0 | 55.5% | 34.1% |
| AutomationBench | 44.7% | 10.7% |
| GDPval-AA v2.1 (Elo) | 1844 | 1449 |
| Chartography (no tools) | 61.6% | 15.6% |
| FrontierCode v1.1 | 46.2% | 42.4% |
The Terminal-Bench line deserves a caveat. The 10.3% for Sonnet 5 appears only in the announcement chart; the system card doesn’t report a Sonnet 5 score for that test. Independent runs back up the direction, though: Vals AI measured Sonnet 5 at 8.08% on its own harness. The same leaderboard reverses the Sonnet vs Opus order, with Opus 5.5 at 61.62% and Sonnet 5.5 at 53.03%, so treat the Terminal-Bench lead over Opus as harness-dependent.
Sonnet 5.5 also has a later knowledge cutoff (June 2026 vs January 2026) and a smaller minimum cacheable prompt (512 tokens instead of 1,024), which helps agents with short system prompts get cache hits.
What we measured: four models, one agent, three runs each
We ran all four models through Atomic Agent v0.6.5, our open-source agent, in cloud mode via OpenRouter, with provider default settings and no effort override. Each model got a fresh, isolated copy of the same small Python project and the same one-line prompt, three times.
The task: a tiny expense-splitting library with five planted bugs (float rounding in money parsing, amounts with three decimals accepted, missing zero padding, leftover cents going to the wrong person, people with a zero balance dropped from the output), a missing settlement algorithm, and a missing command-line tool. The agent had to fix the bugs, implement both features to a written spec, add tests, and get the test suite green. We then scored each result against 32 hidden tests the agent never saw. Cost is the sum of OpenRouter’s billed cost for every request in the run.
| Model | Hidden tests passed | Avg time | Avg steps | Avg cost per run |
|---|---|---|---|---|
| Claude Sonnet 5.5 | 96/96 | 48 s | 5.3 | $0.19 |
| Claude Opus 5.5 | 96/96 | 68 s | 5.7 | $0.33 |
| Claude Sonnet 5 | 96/96 | 123 s | 17.3 | $0.30 |
| GPT-6 Sol | 96/96 | 111 s | 21.0 | $0.52 |

All four models solved the task every time, so this table says nothing about which one is smarter. What it shows is how much work each model does to get there, and that is where the money goes.
Sonnet 5.5 finished in 4 to 6 steps. Sonnet 5 needed 16 to 19: it read files, ran tests, and fixed things one at a time, and each step resends the whole conversation. That matches Anthropic’s line that it “typically needs far fewer tokens to do the same work”, and our 37% saving lands a bit beyond their “up to 30% less per task”.
GPT-6 Sol has the same $2/$10 list price as Sonnet 5.5 and still cost 2.7x more per run, largely because it took about four times as many steps. List price only tells you part of the bill.
Limits of this test. It’s one small, well-specified task at default effort, three runs per model. It rewards models that plan well on a clear spec, which is exactly Sonnet 5.5’s home ground. It doesn’t test long, messy, open-ended work, where Opus 5.5 is supposed to pull ahead. Independent evaluators running at maximum effort got a different cost picture, covered in the next section.
Pricing: same list price, different bill
Sonnet 5.5 costs $2 per million input tokens and $10 per million output, identical to Sonnet 5 and half of Opus 5.5. Whether it’s cheaper per task depends mostly on the effort level you pick: at default settings it uses far fewer tokens than Sonnet 5, but at maximum effort independent tests found it more expensive.
| Model | Input | Output | Cache read | Batch (in/out) |
|---|---|---|---|---|
| Claude Sonnet 5.5 | $2 | $10 | $0.20 | $1 / $5 |
| Claude Sonnet 5 | $2 | $10 | $0.20 | $1 / $5 |
| Claude Opus 5.5 | $4 | $20 | $0.20 | $2 / $10 |
| Claude Haiku 4.5 | $1 | $5 | $0.10 | $0.50 / $2.50 |
| GPT-6 Sol | $2 | $10 | $0.20 | 50% off |
| GPT-6 Astra | $10 | $50 | $1.00 | 50% off |
Prices per million tokens from Anthropic’s pricing page and OpenAI’s pricing page.
The full 1M context window is billed at the standard rate, with no long-context surcharge. US-only inference adds 10%. There is no fast mode for Sonnet 5.5; that option exists only for Opus.
Now the counterweight. Artificial Analysis ran Sonnet 5.5 at max effort and measured $7.60 per task, about 50% more than Sonnet 5, with roughly 193,000 output tokens per task. Vals AI saw the same pattern on Terminal-Bench: $19.33 per task for Sonnet 5.5 against $14.25 for Sonnet 5. At max effort the model thinks a lot, and thinking is billed as output.
The practical rule: don’t run Sonnet 5.5 at max by default. Anthropic’s own docs suggest starting at high on the API and at medium for well-specified agentic coding. The Claude apps and Claude Code already default to medium.
What changed in the API
Sonnet 5.5 is not a drop-in swap for code written against Sonnet 5. Thinking can no longer be switched off, forced tool use now returns an error, and thinking blocks are tied to the account and conversation that produced them. Check the migration guide before switching a production model ID.
The changes most likely to break code written for Sonnet 5:
- Thinking can’t be turned off.
thinking: {"type": "disabled"}returns a 400 error. The lowest setting is the newbetween_toolsmode. - Forced tool use is gone.
tool_choiceset toanyor a specific tool returns a 400. - Thinking blocks are bound to the model, conversation, and account that produced them.
- Computer use on the Claude API and Google Cloud needs the current toolset version.
If you’re coming from Sonnet 4.6 or older, the list is longer: sampling parameters, assistant prefill, and manual budget_tokens all return 400. Sonnet 5.5 has five effort levels (low, medium, high, xhigh, max), with high as the API default.
- Output limit: 128K tokens, or 300K via Message Batches with the
output-300k-2026-03-24beta header.
The model IDs: claude-sonnet-5-5 on the Claude API, Google Cloud, and Microsoft Foundry; anthropic.claude-sonnet-5-5 on Amazon Bedrock; anthropic/claude-sonnet-5.5 on OpenRouter. In Claude Code, the sonnet alias switches to 5.5 from version 2.1.284, while the default model stays Opus 5.5.
The new cyber safeguards, and why your request may go to Sonnet 5
Sonnet 5.5 is the first Sonnet to ship with the cyber safeguards Anthropic reserved for its top models. Anthropic’s reason is that its security capabilities are now “comparable to Opus 5’s”. When a request trips the filter, it can be answered by Sonnet 5 instead, and you’ll see more refusals on security work.
How it works in practice:
- Source code vulnerability research is allowed. Vulnerability discovery in compiled binaries is blocked at general access.
- Fallback to Sonnet 5 is automatic in Anthropic’s apps. On the API it is an opt-in beta (
fallbacks: "default"); without it, you get a refusal. - Reasoning extraction is blocked by a separate classifier with no fallback, and thinking blocks only work in the account that produced them. The Next Web ties this to August research that recovered API keys and passwords from public agent traces.
- A Cyber Verification Program for Sonnet 5.5 is announced but not live yet.
The safeguards also touch the benchmarks. In the Terminal-Bench runs, fallbacks affected 1.5% of Sonnet 5.5 trials but 10% of Opus 5.5 trials, which partly explains why Sonnet edges out Opus on that one test.
Sonnet 5.5 vs GPT-6 Sol and other rivals
GPT-6 Sol is the direct competitor: same $2/$10 price and launched the same month. On the benchmarks both companies report, Sonnet 5.5 leads on office-work and chart tasks, GPT-6 Sol leads on FrontierCode, and in our agent test Sonnet 5.5 was 2.3x faster and 2.7x cheaper on the same task.
| Benchmark | Sonnet 5.5 | GPT-6 Sol |
|---|---|---|
| FrontierCode v1.1 | 46.2% | 49.3% |
| GDPval-AA v2.1 (Elo) | 1844 | 1487 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1483 |
| AutomationBench | 44.7% | 32.0% |
| Chartography (no tools) | 61.6% | 53.6% |
Figures from Anthropic’s announcement and system card; GPT-6 Sol numbers are as reported by Anthropic.
If you want a cheaper tier altogether, the field is wide. GLM-5.3 lists at $1.40/$4.40, Qwen3.8-Max at $2/$6, and GPT-6 Luna at $0.10/$0.50. Anthropic says Haiku 5.5 is coming “in the coming weeks”; until then Haiku 4.5 at $1/$5 is its budget option.
How to try Sonnet 5.5
The fastest way is through the Claude API or Claude Code. If you want to test it against other models on your own tasks, the setup we used above takes a few minutes: Atomic Agent is open source, connects to OpenRouter, and lets you switch the model per run while keeping the same tools and prompts.
curl -fsSL https://atomicagent.io/install | shatomic-agent tui# then pick OpenRouter and anthropic/claude-sonnet-5.5 in /modelsOn Windows, use irm https://atomicagent.io/install.ps1 | iex in PowerShell.
The same agent also runs local models through llama.cpp, so you can put Sonnet 5.5 side by side with a model on your own machine. We wrote up how that trade-off looks in local AI agent vs cloud. If you’re comparing coding agents rather than models, see our list of open-source Claude Code alternatives.
The takeaway
Sonnet 5.5 moves the default. For clear, well-scoped work, it gets within a couple of points of Opus 5.5 on most agentic benchmarks at half the price, and in our test it did the job in a third of Sonnet 5’s steps. Keep Opus 5.5 for large, open-ended coding and hard reasoning, where it still leads, most clearly on SWE-Bench Pro.
Two things to watch: effort level drives your bill more than list price, and security teams should test the new safeguards before switching.
Benchmarks, pricing, and API details are from Anthropic’s Sonnet 5.5 announcement, system card, and platform docs (September 28, 2026). Our test ran on September 29, 2026 with Atomic Agent v0.6.5 via OpenRouter; task spec and per-run numbers are available on request.


