Claude Sonnet 5.5 scores 18 points more than Sonnet 5 and the list price did not move
Sonnet 5.5 scores 56 against Sonnet 5's 38 on the same $2 / $10 list price. Five breaking changes, recalibrated effort levels, and a per-task figure that is not a like-for-like comparison.
GTM Architect & Growth Operator · Now · 1 October 2026 · 7 min read
TL;DR · Key insights
- Sonnet 5.5 scores 56 on Intelligence Index v4.3.2 against Sonnet 5's 38. That is 18 points for the same $2 / $10 list price, the largest jump Anthropic has shipped inside one tier.
- The price genuinely did not move. Sonnet 5's $2 / $10 started as introductory pricing to 31 August 2026, and Anthropic cancelled the rise to $3 / $15 that was scheduled for 1 September.
- Five breaking changes hit code already running on Sonnet 5, one more than the Opus 5.5 port had. Effort levels are also recalibrated, so the setting you tuned does not mean what it meant.
- At 56 for $7.62 per index task, Sonnet 5.5 costs more per finished task than Opus 5.5, which scores 58 for $5.98. On this benchmark the cheaper model per token is the dearer one per job.
Claude Sonnet 5.5 shipped on 28 September 2026. It scores 56 on the Artificial Analysis Intelligence Index where Sonnet 5 scores 38, and the price sheet did not move: $2 per million input tokens, $10 output, same as before.
Eighteen points for nothing is not a trade. It is the part of this release worth reading twice, because every other upgrade in this family over the past year moved the price, or the token count, or quietly moved both while the headline talked about benchmarks.
The price really did not move
Identical down the whole row. $2 input, $10 output, $0.20 cache reads, $2.50 and $4 cache writes.
Worth saying plainly, because Sonnet 5’s price was never meant to hold: the $2 / $10 was introductory pricing through 31 August 2026, with a rise to $3 / $15 scheduled for 1 September that Anthropic then cancelled, turning the introductory number into the standard one. So this is not quite “new model, same price”. It is “new model, and the old discount became permanent on the way”.
Eighteen points is a wide jump inside one tier
Sonnet 5.5 scores 56. Opus 5.5 scores 58, on the same index version, v4.3.2. Sonnet 5 scored 38, a long way below both.
A tier upgrade that closes most of the distance to the flagship, at a quarter of the flagship’s output price, is the kind of change that makes you re-test rather than re-read. Six months ago you might have picked Opus on a margin of a dozen points. That margin is now two.
The per-task figure is not a like-for-like comparison
Artificial Analysis puts Sonnet 5.5 at $7.62 per index task and Sonnet 5 at $5.09. That looks like a 50% increase. I am not going to call it one, because the two runs are not configured the same way and the arithmetic does not support the headline.
Sonnet 5.5 is measured as Adaptive Reasoning, Max Effort, Default Fallback. Sonnet 5 is measured as Adaptive Reasoning, Max Effort, with no fallback, and Artificial Analysis now marks its results historical and no longer updated. The token counts are 410M against 370M, about 11% apart. At identical per-token pricing, 11% more tokens cannot produce 50% more cost. The gap lives in how the runs were set up, not only in how much each model writes.
What you can take from it: Sonnet 5.5 is verbose. 410M tokens against a median of 82M across comparable models is five times the middle of the field. Whatever your per-task bill was, budget for it to go up, and measure it on your own workload rather than borrowing either figure.
The cheap model is not the cheap one here
This is the number that should change a decision.
| Sonnet 5.5 | Opus 5.5 | |
|---|---|---|
| Index score, v4.3.2 | 56 | 58 |
| List price | $2 / $10 | $4 / $20 |
| Cost per index task | $7.62 | $5.98 |
| Tokens on the run | 410M | 260M |
Sonnet 5.5 costs half as much per token and more per finished task. Opus 5.5 scores higher and bills less getting there. The reason is length: it reaches an answer in 260M tokens where Sonnet 5.5 takes 410M.
That inverts the usual reason for picking Sonnet. If you chose it to save money on long agentic runs, check the bill before you keep that reasoning. If you chose it for latency, the logic holds: a cheaper, faster model that writes more is still faster per token, and the index does not measure the thing you care about.
Five things that will break your port
Anthropic documents five breaking changes against code already running on Sonnet 5:
- Up-front thinking is turned off with
between_tools, not the control you used before. - Forced tool use returns an error.
- Thinking blocks are tied to the model and the conversation that produced them.
- On the Claude API and Google Cloud, the earlier
computer_20251124computer use tool is not accepted. - The advisor tool rejects Claude Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.
That is one more than the Opus 5 to Opus 5.5 port had, and items two, three and four are the same three that caught people there. If you did that migration, you have already written most of this one.
Effort levels mean something else now
The quieter change is not on the breaking list. Effort levels are recalibrated: the same setting does not produce the same amount of thinking it did on Sonnet 5. Anthropic’s instruction is to re-run your effort sweep rather than carrying a setting over.
Their starting points: high unless the workload is agentic or latency-sensitive. For agentic coding and multistep tool use, medium for well-specified tasks, high for harder or longer ones. For chat and anything latency-sensitive, medium or low.
If you tuned an effort level against a cost ceiling, that ceiling moved without anyone telling you.
Which one I would use
Sonnet 5.5. That one is quick: Sonnet 5 is on the legacy list, its benchmark results are frozen, and it scores 18 points lower for the same money. There is no version of this where staying put is the better call.
The decision that is actually open is Sonnet 5.5 against Opus 5.5, and it is closer than the price sheet suggests. Two points apart on the index, and the cheaper one bills more per task on that benchmark.
What would change my mind
If your tasks are short, the per-task figures above stop mattering and the list price is what reaches your bill. Then Sonnet 5.5 at half of Opus 5.5 is the obvious pick and the verbosity never gets the chance to cost you anything.
I would also want a second measurement before leaning hard on the $7.62. One run, one configuration, one benchmark that rewards long reasoning is a thin basis for a budget. Run your own tasks, count your own tokens, then decide.
Questions people asked
Is Claude Sonnet 5.5 cheaper than Claude Sonnet 5?
No, and it is not dearer either. Both list at $2 per million input tokens and $10 output, with the same $0.20 cache reads and the same $2.50 and $4 cache writes. Anthropic shipped the upgrade without touching the price, which is unusual: Sonnet 5’s $2 / $10 was introductory pricing through 31 August 2026, and the increase to $3 / $15 scheduled for 1 September never happened.
Is Claude Sonnet 5.5 better than Claude Sonnet 5?
By 18 points on the independent index, which is a wide margin inside one tier. Artificial Analysis scores Sonnet 5.5 at 56 and Sonnet 5 at 38 on Intelligence Index v4.3.2. Anthropic has moved Sonnet 5 to its legacy list and Artificial Analysis marks it deprecated, so Sonnet 5.5 is the Sonnet you get unless you pin the older model ID.
What breaks when I upgrade from Claude Sonnet 5 to Sonnet 5.5?
Five things, per Anthropic’s own migration notes. Up-front thinking is turned off with between_tools rather than the old control. Forced tool use returns an error. Thinking blocks are tied to the model and the conversation that produced them. On the Claude API and Google Cloud the earlier computer_20251124 computer use tool is not accepted. And the advisor tool rejects Claude Opus 4.8, Opus 4.7 and Sonnet 5 as advisors. A sixth change is not a break but will change your output: effort levels are recalibrated, so Anthropic tells you to re-run your effort sweep instead of carrying a setting over.
Should I use Claude Sonnet 5.5 or Claude Opus 5.5?
On list price Sonnet 5.5 is half of Opus 5.5: $2 / $10 against $4 / $20. On the index it is not the cheap option. Artificial Analysis measured $7.62 per task for Sonnet 5.5 at a score of 56, and $5.98 for Opus 5.5 at 58. Sonnet 5.5 wrote 410M tokens on that run to Opus 5.5’s 260M. If your work looks like the index, the per-token saving does not survive contact with the token count. If your tasks are short or latency matters, the list price is the one that reaches your bill.
For the tier above, see Opus 5.5 against Opus 5. Every comparison on this site is collected in one place at the model comparisons.

