Anthropic released Claude Sonnet 5.5 on September 28, 2026, the second model in its Claude 5.5 family and successor to the mid-tier Sonnet 5 from June.[1][4] The pitch is unusual: the list price is unchanged, yet Anthropic says the model generates output more than 30% faster and costs up to 30% less per task, because it needs fewer tokens and fewer tool calls to finish the same job.[1][5] If that holds at production scale, it shifts what teams optimise for - not the price of one token, but the price of a finished task.
Why Cost Per Task Matters More Than Token Price
Sonnet 5.5 keeps Sonnet 5's price list exactly: $2 per million input tokens, $10 per million output tokens, $0.20 per million cache-read tokens and $2.50 per million cache-write tokens, against $4 input, $20 output and $5 cache writes for the larger Opus 5.5.[1][5] The saving therefore comes from doing less work, not from a cheaper meter. Enterprises increasingly evaluate models this way: every unnecessary tool call, failed action and retry adds latency, cost and another chance for an autonomous workflow to go off course.[5]
Benchmarks: The Mid-Tier Is Catching the Flagship
On Anthropic's reported scores, Sonnet 5.5 lands close to Opus 5.5 in several domains while comfortably beating its own predecessor:[1]
- Agentic coding (Terminal-Bench 4.0): 70.6% for Sonnet 5.5, against 10.3% for Sonnet 5 and 66.4% for Opus 5.5.
- Real-world knowledge work (GDPval-AA v2.1): 1844, against 1449 for Sonnet 5 and 1846 for Opus 5.5.
- Long-horizon knowledge work (AA-Briefcase v1.1): 1811, versus 1359 for Sonnet 5 and 1822 for Opus 5.5.
- Computer use (OSWorld 2.1, partial credit): 80.1%, against 57.0% for Sonnet 5 and 81.8% for Opus 5.5.
- Visual chart recognition (Chartography): 61.6%, versus 15.6% for Sonnet 5 and 64.4% for Opus 5.5.
- Coding agents on real sessions (CursorBench 4.0): 55.5%, against 34.1% for Sonnet 5 and 57.8% for Opus 5.5.
Anthropic is candid about the limits of its own table: benchmark scores capture only one facet of a model's capabilities, and by the company's own account Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgement.[1] A footnote notes that Artificial Analysis measured Sonnet 5.5 on a pre-release deployment carrying a structured-outputs bug, which would if anything understate its scores.[1]
What Early Testers Reported
Anthropic's launch post relies on companies that tested the model before release. These are first-party accounts, not independent benchmarks, but they show a consistent pattern of fewer steps and less token burn:
- Box reported results 2.4 times faster with 12% fewer total tokens, and said the model re-checked source documents and caught errors Sonnet 5 missed.[1]
- Zendesk processed support tickets 20% faster while making fewer incorrect escalation decisions than the Claude models it runs in production.[1]
- Slack measured roughly 14% fewer output tokens on its offline Slackbot evaluations without changing any prompts.[1]
- Lovable's coding evaluations showed about a third fewer tool calls and roughly half the shell executions needed to finish a task.[1]
- Base44, building 118 real applications, saw Sonnet 5.5 match Opus 5 quality in an average of 3.6 iterations per build where Opus 5 needed 7.7, with the fewest failed tool calls of the models compared.[1]
- Balyasny Asset Management used about 121,000 tokens per answer on a private suite of 2,441 finance tasks, where Sonnet 5 burned roughly 497,000.[1]
Claude Sonnet 5.5 shows better judgment than Sonnet 5 across different levels of complexity, while spending significantly fewer output tokens. The tendency to reach for web search too often and the high token use are both gone in this new model.[1]
Breaking Changes Developers Must Handle
Sonnet 5.5 is not a drop-in replacement. The model identifier is claude-sonnet-5-5, and five breaking changes affect code already running on Sonnet 5.[2][3]
- Adaptive thinking now runs by default. Sending
thinking: {"type": "disabled"}returns a 400 error; the replacement isbetween_tools, the lowest thinking setting, accepted at high effort and below.[2][3] - Forced tool choice is gone: the
tool_choicevaluestoolandanyare rejected, so sendautoand mark toolsstrict: trueinstead, up to 20 strict tools per request.[3] - Thinking blocks are signed over the conversation and tied to the account that produced them, so transcripts must stay append-only. Replaying one after editing earlier history returns a 400 error on accounts created on or after 31 August 2026.[3]
- On the Claude API and Google Cloud, computer use requires the
computer_toolset_20260801toolset; the oldercomputer_20251124version is refused.[3] - The minimum cacheable prompt drops to 512 tokens, and longer remarks between tool calls now arrive in thinking blocks rather than text blocks, so an interface that streamed that text goes quiet until it reads the new block type or turns up-front thinking off.[2][3]
Effort levels have also been recalibrated to low, medium, high, xhigh and max. Claude Code and the Claude apps default to medium, while the Claude Platform defaults to high, buying a more thorough pass over the work at the cost of latency and tokens.[1][2][3]
Cyber Safeguards Reach the Sonnet Tier
The capability jump brings a policy change. Anthropic says Sonnet 5.5's cybersecurity abilities are comparable to Opus 5's, making it the first Sonnet model to ship with the same class of cyber safeguards and fallbacks used on the company's most capable systems.[1] In practice, higher-risk security requests visibly fall back to Sonnet 5, while routine software development and vulnerability remediation keep working normally.[1][5][6] Refusals are categorised as cyber, bio, frontier_llm, reasoning_extraction or general_harms, and only cyber and frontier_llm declines are retried server-side on Sonnet 5.[3]
Sonnet 5.5 is also the first Sonnet model to launch with classifiers meant to block distillation attacks, in which attackers use fake accounts at industrial scale to copy a model's capabilities, and it extends preserved thinking so reasoning cannot be decoupled from the account that produced it.[1] For engineering teams, the practical consequence is worth testing early: the model answering a request may not be the one you asked for.[6] Anthropic says it is still tuning the classifiers to reduce false positives and plans to widen its Cyber Verification Program so vetted defenders get more advanced capabilities with fewer restrictions.[6]
Availability and What Comes Next
Sonnet 5.5 is available now through the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry with zero data retention, at Sonnet 5's price, with a 1M-token context window and a 128K maximum output.[1][2] Claude Haiku 5.5, aimed at high-volume and cost-sensitive workloads, follows in the coming weeks.[1][4] The competitive backdrop is tight: OpenAI's GPT-6 Sol lists at the same $2 and $10 per million tokens, while Google's Gemini 3.8 Flash undercuts both with introductory rates through the end of 2026.[5] Anthropic's counter-argument: a model finishing a job in fewer reasoning steps and tool calls can be cheaper to run even when the per-token rate does not move.[5]
Sources
- Introducing Claude Sonnet 5.5 (Anthropic, 28 Sep 2026)
- Claude Sonnet 5.5 model overview (Claude Platform Docs)
- Migrating to Claude Sonnet 5.5 (Claude Platform Docs)
- Anthropic releases Sonnet 5.5 (TechCrunch, 28 Sep 2026)
- Anthropic launches Claude Sonnet 5.5 (VentureBeat, 28 Sep 2026)
- You picked Claude Sonnet 5.5 - but Anthropic may send your request to Sonnet 5 (The New Stack, 28 Sep 2026)
admin
Comments