TL;DR
Anthropic has released Sonnet 5.5, the second member of its Claude 5.5 family after Opus 5.5. List prices are unchanged from Sonnet 5, but Anthropic says it gets through the same work on fewer tokens, cutting per-task cost by as much as 30%. On several of Anthropic’s benchmarks it lands close to Opus 5.5, which costs twice as much per token, so for many everyday business workloads the cheaper model is worth testing as the default.
Where it lands
The standout number is agentic coding. On Terminal-Bench 4.0, Sonnet 5.5 scores 70.6% against Sonnet 5’s 10.3%, and above the 66.4% Opus 5.5 reached at its highest effort setting. GDPval-AA, a test of practical work spanning 44 occupations, puts it two points behind Opus 5.5, and roughly 400 points clear of the previous Sonnet. In computer use, it scores 80.1% on OSWorld 2.1 against 81.8% for Opus.
Anthropic is careful not to oversell it. Opus 5.5, it says, remains “clearly stronger at complex, open-ended work requiring sustained judgment”. Sonnet is pitched at bounded, day-to-day jobs: bug fixes, documents, slides and spreadsheets.
Price and speed
Input stays at $2 per million tokens and output at $10, half the rate for Opus 5.5 on both. Cache reads cost $0.20 for either model. Anthropic calls it its quickest Sonnet yet, with output arriving 30%+ sooner than from Sonnet 5. Effort settings matter too: the apps and Claude Code start at Medium, the Claude Platform API at High. On some benchmarks, Anthropic says, Sonnet 5.5 at Low or Medium beats the best Sonnet 5 result for roughly 10% of the per-task cost.
Safety changes
Because its cyber abilities now rival Opus 5, Sonnet 5.5 is the first Sonnet to ship with cyber safeguards; riskier security requests are handed back to Sonnet 5, and users can see that happen. It is also the first with classifiers that block attempts to extract its reasoning at scale, a tactic known as distillation. On Anthropic’s newer containment tests it is, the company says, the least likely of its models to probe the edges of its sandbox, a measure that carries more weight after this week’s AISI findings on agent behaviour.
Looking forward
Haiku 5.5 should follow within weeks. Sonnet 5.5 is available now on the Claude Platform and through the big three clouds (AWS, Azure and Google Cloud), with zero data retention on offer. For UK teams running Opus 5.5 on routine work, testing whether Sonnet handles the same tasks at half the token price is an easy win. Benchmark gaps rarely survive contact with a specific workflow, so it is worth checking on your own tasks.