Sonnet 5.5 vs Opus 5.5: the per-task math that changes your bill
Sonnet 5.5 vs Opus 5.5: the per-task math, the effort dial, caching and when not to migrate. Anthropic's numbers, with caveats.
Sonnet 5.5 vs Opus 5.5 is the comparison that matters since September 28, 2026, when Anthropic launched Claude Sonnet 5.5 at the same list price as Sonnet 5: $2 per million input tokens and $10 per million output tokens. What changed is how much work it does for that money, and the practical rule is to use Sonnet by default and Opus only when Sonnet fails on your own tests.
What changed with Sonnet 5.5 versus Sonnet 5
Anthropic says the model responds over 30% faster and finishes a task spending up to 30% less. It has a one-million-token context window and is available in the Claude apps, Claude Code and on Amazon, Google and Microsoft clouds. Note that the 30% is a ceiling, not a promised average.

Price per token versus cost per task
We compare models by price per token, but nobody pays for tokens: you pay for tasks. It's like a car: the price per liter says nothing if you don't know how many kilometers it gets. Balyasny, an investment firm, ran 2,441 finance tasks. With the previous model each answer used 497,000 tokens; with Sonnet 5.5, 121,000.
In a simplified calculation, everything at input price, that's almost a dollar per answer versus 24 cents. Multiplied by the 2,441 tasks: $2,426 versus $591. A Sapio Research survey for DoiT of 500 finance leaders found 79% had AI cost overruns in the last twelve months, and only 15% can calculate their return without major bottlenecks.
Sonnet 5.5 vs Opus 5.5: what the tests say
Versus Sonnet 5 the jump is large: on Terminal-Bench 4.0 it goes from 10% to 70%, and on the OSWorld computer-use test, from 57% to 80%. Versus Opus 5.5, which costs double per token, it lands very close on office work: 1844 points against 1846. On Terminal-Bench it even beats it, 70.6 to 66.4, though Opus was measured at extra-high effort.
So why use Opus at all? According to Anthropic, it's still clearly better at open-ended work that demands sustained judgment: on FrontierCode it scores 54.4 against 46.2. If your task has a clear scope, paying double for a two-point gap on office work is paying for something you don't use.
The three levers on your bill
The effort dial. There are five levels: low, medium, high, extra high and max. Anthropic claims that on Terminal-Bench at medium effort it beats the previous model's best result at under a tenth of the cost per task. The levels were recalibrated, so your old setting no longer means the same. The official advice is to start at high, or medium for agents with well-defined tasks, and measure.
Caching. An agent that loads 40,000 tokens of manual on every query, a thousand times a day, spends about $80 without caching. With caching, reading them costs 20 cents per million: about $8 plus 10 cents to write it once. That's our own estimate with official prices, and it depends on the cache staying alive between queries.
Batches. If your work can wait, the batch interface charges half.
How to migrate your app
Change the model name, one line, and run your tests. If you sent temperature or top-p with a non-default value, the API returns a 400 error. If you forced tool use, that also fails: now you ask for it in the prompt and leave automatic mode. If you turned thinking off with the disabled value, you now use between_tools.
What can go wrong
- If your app shows users the text the model writes between tool calls, it may go silent with no error: that text now arrives in thinking blocks, empty by default.
- Effort levels changed, and inherited settings may think more or less than you expect.
- Thinking blocks are tied to the model and account that produced them: if you switch to Opus mid-conversation, that reasoning is lost.
- On higher-risk cybersecurity tasks, the system visibly falls back to Sonnet 5.
How to apply it in your business?
Pick 50 real tasks with a clear success criterion, run them on your current model and on Sonnet 5.5 at medium and high effort, and record tokens, steps and successes. Calculate cost per completed task, counting retries, and decide with your own numbers. Ideal cases are batch documents (Zendesk reported tickets processed 20% faster; Box, 2.4 times faster and 12% fewer tokens) and agencies building apps (Base44 measured 3.6 iterations on average versus 7.7 for Opus 5). That last figure compares against Opus 5, not 5.5, and it's the customer's measurement.
For the quick version, read the full breakdown in Short format of Claude Sonnet 5.5.
Frequently asked questions
How much does Sonnet 5.5 cost?
$2 per million input tokens and $10 per million output tokens, the same as Sonnet 5.
Is Sonnet 5.5 better than Opus 5.5?
Not at everything. It nearly ties on office work (1844 vs 1846), but Opus is still better at open-ended work that demands sustained judgment.
When should I use Opus 5.5?
When Sonnet 5.5 fails on your own tests, not when it sounds more impressive.
Are the numbers independent?
No. Almost all come from Anthropic and its own customers, and none has been repeated by third parties.
Conclusion
In Sonnet 5.5 vs Opus 5.5, the list price didn't change but the bill did: fewer tokens, fewer steps, adjustable effort and caching. It all shows up in the bill, but only if you measure it. Don't compare price per token: compare cost per task.