Small Claude character sprinting ahead of a bigger one carrying a sack of coins and an older one far behind, Atomic Agent mascot at the finish with a stopwatch
Models

Claude Sonnet 5.5 vs Opus 5.5 vs Sonnet 5: Benchmarks, Pricing, and Which to Use

Published
Reading time
11 min read
Tag

Sonnet 5.5 vs Opus 5.5 vs Sonnet 5: full benchmarks, pricing, and our own test. Same $2/$10 price, 37% cheaper per task and 2.6x faster in our runs.

Anthropic released Claude Sonnet 5.5 on September 28, 2026, six days after Opus 5.5. The API model ID is claude-sonnet-5-5. It keeps Sonnet 5’s price of $2 per million input tokens and $10 per million output, and Anthropic says it generates output 30%+ faster and costs up to 30% less per task because it needs fewer tokens.

We didn’t want to repeat the launch post, so we ran it. Same coding task, same agent, three runs each for Sonnet 5.5, Sonnet 5, Opus 5.5, and OpenAI’s GPT-6 Sol, the model Anthropic itself compares against.


Sonnet 5.5 vs Opus 5.5: which should you use?

Use Sonnet 5.5 when the task has a clear spec: bug fixes, well-defined features, documents, repeatable agent jobs. Use Opus 5.5 for long, open-ended work that needs judgment across many steps. Opus 5.5 costs exactly twice as much per token ($4/$20) and still leads on repository-scale coding, most clearly on SWE-Bench Pro (89.9% vs 81.3%).

Anthropic’s own framing matches that split. It calls Sonnet 5.5 “a faster, lower-cost complement to Claude Opus 5.5” and says it is “strongest at well-scoped everyday tasks, fixing bugs,” documents, slides, and spreadsheets. Its prompting guide is blunter: “for the hardest long-horizon work, an Opus model is the better choice.”

Here is the gap, benchmark by benchmark:

BenchmarkSonnet 5.5Opus 5.5Gap
SWE-Bench Pro81.3%89.9%Opus +8.6
FrontierCode v1.146.2%54.4%Opus +8.2
Humanity’s Last Exam (with tools)64.5%67.7%Opus +3.2
CursorBench 4.055.5%57.8%Opus +2.3
OSWorld 2.180.1%81.8%Opus +1.7
GDPval-AA v2.1 (Elo)18441846tie
AutomationBench44.7%42.5%Sonnet +2.2
Terminal-Bench 4.070.6%66.4%Sonnet +4.2

Anthropic’s announcement figures. Sonnet 5.5 is shown at max effort; on FrontierCode it scores higher at xhigh (52.1%), which narrows that gap to about 2 points. Opus 5.5’s Terminal-Bench score is at xhigh (64.8% at max).

Opus keeps a clear lead on SWE-Bench Pro and a smaller one on hard reasoning. On computer use, office work, and terminal tasks, Sonnet 5.5 sits within two points or ahead, which is the ground most agents actually cover.

In our run the two models behaved almost identically: about 6 steps each, every hidden test passed, but Opus 5.5 took 68 seconds and $0.33 per run against Sonnet 5.5’s 48 seconds and $0.19. On a task this size, Opus bought nothing extra.

Sonnet 5.5 vs Sonnet 5: is the upgrade worth it?

Yes. The per-token price is identical, Sonnet 5.5 scores higher on every benchmark in Anthropic’s table, and in our test it did the same job 2.6x faster for 37% less money. The only reasons to wait are the new cyber safeguards and a handful of API changes that break old request code, both covered below.

The jumps are large for a point release:

BenchmarkSonnet 5.5Sonnet 5
Terminal-Bench 4.070.6%10.3%
OSWorld 2.1 (partial score)80.1%57.0%
SWE-Bench Pro81.3%63.2%
CursorBench 4.055.5%34.1%
AutomationBench44.7%10.7%
GDPval-AA v2.1 (Elo)18441449
Chartography (no tools)61.6%15.6%
FrontierCode v1.146.2%42.4%

The Terminal-Bench line deserves a caveat. The 10.3% for Sonnet 5 appears only in the announcement chart; the system card doesn’t report a Sonnet 5 score for that test. Independent runs back up the direction, though: Vals AI measured Sonnet 5 at 8.08% on its own harness. The same leaderboard reverses the Sonnet vs Opus order, with Opus 5.5 at 61.62% and Sonnet 5.5 at 53.03%, so treat the Terminal-Bench lead over Opus as harness-dependent.

Sonnet 5.5 also has a later knowledge cutoff (June 2026 vs January 2026) and a smaller minimum cacheable prompt (512 tokens instead of 1,024), which helps agents with short system prompts get cache hits.

What we measured: four models, one agent, three runs each

We ran all four models through Atomic Agent v0.6.5, our open-source agent, in cloud mode via OpenRouter, with provider default settings and no effort override. Each model got a fresh, isolated copy of the same small Python project and the same one-line prompt, three times.

The task: a tiny expense-splitting library with five planted bugs (float rounding in money parsing, amounts with three decimals accepted, missing zero padding, leftover cents going to the wrong person, people with a zero balance dropped from the output), a missing settlement algorithm, and a missing command-line tool. The agent had to fix the bugs, implement both features to a written spec, add tests, and get the test suite green. We then scored each result against 32 hidden tests the agent never saw. Cost is the sum of OpenRouter’s billed cost for every request in the run.

ModelHidden tests passedAvg timeAvg stepsAvg cost per run
Claude Sonnet 5.596/9648 s5.3$0.19
Claude Opus 5.596/9668 s5.7$0.33
Claude Sonnet 596/96123 s17.3$0.30
GPT-6 Sol96/96111 s21.0$0.52

Bar charts of average cost and time per run: Sonnet 5.5 $0.19 and 48 s, Opus 5.5 $0.33 and 68 s, Sonnet 5 $0.30 and 123 s, GPT-6 Sol $0.52 and 111 s

All four models solved the task every time, so this table says nothing about which one is smarter. What it shows is how much work each model does to get there, and that is where the money goes.

Sonnet 5.5 finished in 4 to 6 steps. Sonnet 5 needed 16 to 19: it read files, ran tests, and fixed things one at a time, and each step resends the whole conversation. That matches Anthropic’s line that it “typically needs far fewer tokens to do the same work”, and our 37% saving lands a bit beyond their “up to 30% less per task”.

GPT-6 Sol has the same $2/$10 list price as Sonnet 5.5 and still cost 2.7x more per run, largely because it took about four times as many steps. List price only tells you part of the bill.

Limits of this test. It’s one small, well-specified task at default effort, three runs per model. It rewards models that plan well on a clear spec, which is exactly Sonnet 5.5’s home ground. It doesn’t test long, messy, open-ended work, where Opus 5.5 is supposed to pull ahead. Independent evaluators running at maximum effort got a different cost picture, covered in the next section.

Pricing: same list price, different bill

Sonnet 5.5 costs $2 per million input tokens and $10 per million output, identical to Sonnet 5 and half of Opus 5.5. Whether it’s cheaper per task depends mostly on the effort level you pick: at default settings it uses far fewer tokens than Sonnet 5, but at maximum effort independent tests found it more expensive.

ModelInputOutputCache readBatch (in/out)
Claude Sonnet 5.5$2$10$0.20$1 / $5
Claude Sonnet 5$2$10$0.20$1 / $5
Claude Opus 5.5$4$20$0.20$2 / $10
Claude Haiku 4.5$1$5$0.10$0.50 / $2.50
GPT-6 Sol$2$10$0.2050% off
GPT-6 Astra$10$50$1.0050% off

Prices per million tokens from Anthropic’s pricing page and OpenAI’s pricing page.

The full 1M context window is billed at the standard rate, with no long-context surcharge. US-only inference adds 10%. There is no fast mode for Sonnet 5.5; that option exists only for Opus.

Now the counterweight. Artificial Analysis ran Sonnet 5.5 at max effort and measured $7.60 per task, about 50% more than Sonnet 5, with roughly 193,000 output tokens per task. Vals AI saw the same pattern on Terminal-Bench: $19.33 per task for Sonnet 5.5 against $14.25 for Sonnet 5. At max effort the model thinks a lot, and thinking is billed as output.

The practical rule: don’t run Sonnet 5.5 at max by default. Anthropic’s own docs suggest starting at high on the API and at medium for well-specified agentic coding. The Claude apps and Claude Code already default to medium.

What changed in the API

Sonnet 5.5 is not a drop-in swap for code written against Sonnet 5. Thinking can no longer be switched off, forced tool use now returns an error, and thinking blocks are tied to the account and conversation that produced them. Check the migration guide before switching a production model ID.

The changes most likely to break code written for Sonnet 5:

  • Thinking can’t be turned off. thinking: {"type": "disabled"} returns a 400 error. The lowest setting is the new between_tools mode.
  • Forced tool use is gone. tool_choice set to any or a specific tool returns a 400.
  • Thinking blocks are bound to the model, conversation, and account that produced them.
  • Computer use on the Claude API and Google Cloud needs the current toolset version.

If you’re coming from Sonnet 4.6 or older, the list is longer: sampling parameters, assistant prefill, and manual budget_tokens all return 400. Sonnet 5.5 has five effort levels (low, medium, high, xhigh, max), with high as the API default.

  • Output limit: 128K tokens, or 300K via Message Batches with the output-300k-2026-03-24 beta header.

The model IDs: claude-sonnet-5-5 on the Claude API, Google Cloud, and Microsoft Foundry; anthropic.claude-sonnet-5-5 on Amazon Bedrock; anthropic/claude-sonnet-5.5 on OpenRouter. In Claude Code, the sonnet alias switches to 5.5 from version 2.1.284, while the default model stays Opus 5.5.

The new cyber safeguards, and why your request may go to Sonnet 5

Sonnet 5.5 is the first Sonnet to ship with the cyber safeguards Anthropic reserved for its top models. Anthropic’s reason is that its security capabilities are now “comparable to Opus 5’s”. When a request trips the filter, it can be answered by Sonnet 5 instead, and you’ll see more refusals on security work.

How it works in practice:

  • Source code vulnerability research is allowed. Vulnerability discovery in compiled binaries is blocked at general access.
  • Fallback to Sonnet 5 is automatic in Anthropic’s apps. On the API it is an opt-in beta (fallbacks: "default"); without it, you get a refusal.
  • Reasoning extraction is blocked by a separate classifier with no fallback, and thinking blocks only work in the account that produced them. The Next Web ties this to August research that recovered API keys and passwords from public agent traces.
  • A Cyber Verification Program for Sonnet 5.5 is announced but not live yet.

The safeguards also touch the benchmarks. In the Terminal-Bench runs, fallbacks affected 1.5% of Sonnet 5.5 trials but 10% of Opus 5.5 trials, which partly explains why Sonnet edges out Opus on that one test.

Sonnet 5.5 vs GPT-6 Sol and other rivals

GPT-6 Sol is the direct competitor: same $2/$10 price and launched the same month. On the benchmarks both companies report, Sonnet 5.5 leads on office-work and chart tasks, GPT-6 Sol leads on FrontierCode, and in our agent test Sonnet 5.5 was 2.3x faster and 2.7x cheaper on the same task.

BenchmarkSonnet 5.5GPT-6 Sol
FrontierCode v1.146.2%49.3%
GDPval-AA v2.1 (Elo)18441487
AA-Briefcase v1.1 (Elo)18111483
AutomationBench44.7%32.0%
Chartography (no tools)61.6%53.6%

Figures from Anthropic’s announcement and system card; GPT-6 Sol numbers are as reported by Anthropic.

If you want a cheaper tier altogether, the field is wide. GLM-5.3 lists at $1.40/$4.40, Qwen3.8-Max at $2/$6, and GPT-6 Luna at $0.10/$0.50. Anthropic says Haiku 5.5 is coming “in the coming weeks”; until then Haiku 4.5 at $1/$5 is its budget option.

How to try Sonnet 5.5

The fastest way is through the Claude API or Claude Code. If you want to test it against other models on your own tasks, the setup we used above takes a few minutes: Atomic Agent is open source, connects to OpenRouter, and lets you switch the model per run while keeping the same tools and prompts.

Terminal window
curl -fsSL https://atomicagent.io/install | sh
atomic-agent tui
# then pick OpenRouter and anthropic/claude-sonnet-5.5 in /models

On Windows, use irm https://atomicagent.io/install.ps1 | iex in PowerShell.

The same agent also runs local models through llama.cpp, so you can put Sonnet 5.5 side by side with a model on your own machine. We wrote up how that trade-off looks in local AI agent vs cloud. If you’re comparing coding agents rather than models, see our list of open-source Claude Code alternatives.

The takeaway

Sonnet 5.5 moves the default. For clear, well-scoped work, it gets within a couple of points of Opus 5.5 on most agentic benchmarks at half the price, and in our test it did the job in a third of Sonnet 5’s steps. Keep Opus 5.5 for large, open-ended coding and hard reasoning, where it still leads, most clearly on SWE-Bench Pro.

Two things to watch: effort level drives your bill more than list price, and security teams should test the new safeguards before switching.

FAQ

Claude Sonnet 5.5 at a glance.

  • Is Claude Sonnet 5.5 better than Opus 5.5?

    No, not overall. Opus 5.5 still leads on most of Anthropic's benchmarks, including SWE-Bench Pro (89.9% vs 81.3%) and FrontierCode (54.4% vs 46.2%). Sonnet 5.5 comes within about two points on OSWorld, GDPval-AA, and CursorBench at half the list price, and it edges ahead on Terminal-Bench 4.0 and AutomationBench in Anthropic's own runs.

  • How much does Claude Sonnet 5.5 cost?

    Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads are $0.20 per million, 5-minute cache writes $2.50, and the Batch API halves the price to $1 and $5. Opus 5.5 is exactly double at $4 and $20.

  • Is Sonnet 5.5 worth upgrading to from Sonnet 5?

    Yes, for most work. The price per token is identical and Sonnet 5.5 scores higher on every benchmark Anthropic published. In our own small test (one coding task, three runs) it finished 2.6x faster and 37% cheaper. Two things to check first: it has stricter cybersecurity safeguards and a couple of API settings now return errors.

  • What is the context window of Claude Sonnet 5.5?

    Sonnet 5.5 has a 1 million token context window at standard pricing, with no long-context surcharge. Maximum output is 128,000 tokens, or up to 300,000 through the Message Batches API with a beta header. Its knowledge cutoff is June 2026, five months later than Sonnet 5.

  • What is the Claude Sonnet 5.5 model ID?

    The API model ID is claude-sonnet-5-5 on the Claude API, Google Cloud, and Microsoft Foundry. On Amazon Bedrock it is anthropic.claude-sonnet-5-5, with a global inference profile. On OpenRouter it is anthropic/claude-sonnet-5.5. In Claude Code, the sonnet alias points to it from version 2.1.284.

  • Is Sonnet 5.5 cheaper than GPT-6 Sol?

    On list price they are equal: both cost $2 per million input tokens and $10 per million output. Per task it depends on the workload. In our coding test through the same agent, GPT-6 Sol took about four times as many steps and cost $0.52 per run on average, against $0.19 for Sonnet 5.5.


Benchmarks, pricing, and API details are from Anthropic’s Sonnet 5.5 announcement, system card, and platform docs (September 28, 2026). Our test ran on September 29, 2026 with Atomic Agent v0.6.5 via OpenRouter; task spec and per-run numbers are available on request.

Share
Written by
Nadya Dudka Product, Atomic Agent
Andrew Dyuzhov SEO and growth

Run your local agent
in one click