Claude Sonnet 5.5: Faster Everyday Coding Beside Opus 5.5
Anthropic's Claude Sonnet 5.5: faster, cheaper complement to Opus 5.5. Self-reported Terminal-Bench 70.6%, CursorBench 55.5%, FrontierCode Xhigh 52.1%; SWE-bench Pro 81.3%; OSWorld 2.1 strict 43.5% vs partial 80.1%. Same $2/$10 list as Sonnet 5.

Key Numbers
Sonnet 5.5 · Sep 28, 2026
Efficiency & ID
faster · cheaper complement
Anthropic released Claude Sonnet 5.5 on September 28, 2026 as the second model in the Claude 5.5 family, after Opus 5.5. The positioning is a complement, not a crown: faster and cheaper for well-scoped everyday work, while Opus keeps the complex open-ended lane.
That buyer story matters more than any single bench row. Sonnet 5.5 is strongest on everyday coding, bug fixes, and polished docs, slides, and spreadsheets (Anthropic highlights a strong design eye). It is a clear upgrade over Sonnet 5 on coding and knowledge work, with the same $2 / $10 list price and a claim of fewer tokens per task.
Headline self-reported numbers from the announcement table: 70.6% on Terminal-Bench 4.0, 55.5% on CursorBench 4.0, 52.1% on FrontierCode 1.1 Main at Xhigh, 64.5% on HLE with tools. List price is $2 / $10 with $0.20 cache reads. API id: claude-sonnet-5-5. None of the launch scores are LLM Stats verified yet.
At a Glance
- Catalog / API id:
claude-sonnet-5-5only - Organization: Anthropic
- Release: September 28, 2026. Second Claude 5.5 family model after Opus 5.5. Haiku 5.5 still ahead.
- Pricing: $2 input / $0.20 cache read / $2.50 (5m) or $4 (1h) cache write / $10 output per 1M. Batch 50% off. Same list as Sonnet 5.
- Context: 1M input / 128K max output
- Modalities: Text + image in, text out
- Thinking: Adaptive, on by default. Default effort
highon Platform / API;mediumin Claude apps / Claude Code - Knowledge cutoff: June 2026
- Safeguards: First Sonnet with cyber safeguards comparable to Opus-class; biology safeguards same as Sonnet 5
- Secondary card depth: SWE-bench Pro 81.3%, DeepSWE v1.1 71.0%; OSWorld 2.1 strict 43.5% vs partial 80.1%
What's New
Sonnet 5.5 is Anthropic's mid-tier follow-through after Opus 5.5. The product claim is speed and cost for the work most teams actually run every day: scoped coding tasks, bug fixes, and polished office artifacts. Anthropic reports outputs generate 30%+ faster than Sonnet 5, with typically fewer tokens and up to 30% less cost per task at the same list price.
Capability-wise, the announcement table shows large lifts over Sonnet 5 on agentic coding and knowledge-work Elo, and CursorBench within about two points of Opus 5.5. That is the "close enough for most jobs" story, not a claim that Sonnet replaces Opus.
On safeguards, Anthropic says this is the first Sonnet with cyber safeguards comparable to Opus-class models. Cyber capabilities jumped; biology safeguards stay at the Sonnet 5 level. This post does not reprint safety, refusal, or jailbreak tables.
Benchmarks
Prefer Anthropic's announcement table for the hero read. Scores below are self-reported. Not LLM Stats verified. Interpret by job, and keep the Terminal-Bench honesty flag in view before you rank Sonnet above Opus.
Sonnet 5.5 vs peers · percent
Opus at Xhigh · Sonnet may be Max · do not treat as matched effort
primary Xhigh; Max is 46.2% (timeouts / out-of-scope)
suite 2.1 · not 2.0
Coding and agents
Terminal-Bench 4.0 at 70.6%is the loudest coding headline versus Sonnet 5's 10.3% and Opus 5.5's 66.4% (Opus at Xhigh). Treat the Sonnet-vs-Opus margin carefully: third-party and Anthropic notes flag a possible Max-versus-Xhigh mismatch. At matched Xhigh, Anthropic's cost framing still places Sonnet lower. Do not claim Sonnet is overall stronger than Opus. Anthropic says Opus remains clearly stronger at complex open-ended work.
FrontierCode 1.1 Main at 52.1% (Xhigh) is the primary row in this post, versus Sonnet 5 at 42.4% and Opus 5.5 at 54.4%. Max on Sonnet is 46.2%, lower than Xhigh because code-review subagents hit timeouts and out-of-scope edits. CursorBench 4.0 at 55.5%lands within about two points of Opus 5.5's 57.8%, and well above Sonnet 5's 34.1%.
The System Card adds software-engineering depth behind that agent table: SWE-bench Pro 81.3%, Multilingual 90.3%, Multimodal 54.3%, DeepSWE v1.1 71.0%, FrontierSWE V2 61.9% (Proximal harness, mean across trials), FrontierCode Extended 64.4% at Xhigh, and ProgramBench 79.7%. Terminal-Bench-Science 0.1 is 59.9% on the card. Those are absolute card rows, not announcement vs-prior deltas.
Knowledge work and computer use
Knowledge work · Elo
GDPval-AA v2.1
Sonnet 5.5 · announcement table
AA-Briefcase v1.1
Sonnet 5.5 · announcement table
Opus 5.5 · GDPval-AA
same announcement table · within 2 Elo
Sonnet 5 sits at 1449 GDPval-AA and 1359 AA-Briefcase on the same table. Opus 5.5 is 1846 / 1822. Sonnet 5.5 lands within a few Elo of Opus on both rows while keeping the Sonnet sticker.
GDPval-AA v2.1 at 1844 Elo and AA-Briefcase v1.1 at 1811 Elo are the announcement knowledge-work rows (bare titles only). Both sit within a few Elo of Opus 5.5 (1846 / 1822) and far above Sonnet 5 (1449 / 1359). Elo is not a percent bench, so it stays out of the launch bar chart.
HLE with tools at 64.5%improves on Sonnet 5's 54.9% and trails Opus 5.5's 67.7%. The card also reports HLE without tools at 56.9%. Computer use needs both OSWorld 2.1 numbers: the announcement partial score is 80.1%; the System Card strict score is 43.5% on the same suite. That gap is the honest read. Chartography without tools is 61.6% on the launch table; with tools the card is 90.2%.
System card domains
Below are absolute System Card §8 capability rows, grouped by job. Safety, refusal, and jailbreak rows are skipped. No invented peer deltas. Self-reported; not LLM Stats verified.
System card §8 · by job
absolute · no peer rows
Software engineering
Depth behind the agentic coding table. Main FrontierCode stays 52.1% Xhigh.
Proximal harness · mean across trials
Xhigh; Main Xhigh 52.1% / Max Main 46.2%
Computer use
Same OSWorld 2.1 suite as the announcement partial row.
partial on launch table is 80.1%
Vision / CAD
Tool access moves Chartography and BenchCAD a lot.
with tools 90.2%
voxel IoU
Knowledge / office / tools
Office and tool-use depth. Legal all-pass is the hard metric.
all-pass; card mean criterion-pass 92.1%
Health
Prefer Table 8.1.A length-adjusted Professional when citing one number.
length-adjusted Table 8.1.A; raw 77.1%
100 physician EHR tasks · max effort
LA across efforts is a 64.7-65.4% band only
Math
ArXivMath with and without tools.
Multilingual
Broad MMLU-style coverage across languages.
42 languages
11 languages
Bio capability
Capability and knowledge only, including System Card §8.17. Not refusal or jailbreak rows.
Axiom Bio
504 questions
Benchling
Vision and CAD: BenchCAD is 74.7% without tools and 96.3% with a Python tool. Knowledge and office: OfficeQA 76.9% / OfficeQA Pro 65.6%; Legal Agent Benchmark all-pass 11.7% (card mean criterion-pass 92.1%; prefer all-pass as the standard metric); Toolathlon Verified Pass@1 77.8%. AutomationBench is 44.7% (Zapier, max effort, API default fallbacks enabled). Health: prefer length-adjusted HealthBench Professional 69.2% from Table 8.1.A (raw 77.1%); PhysicianBench is 63.2% (100 physician EHR tasks, max effort); HealthBench is stored as raw 69.4% with an LA band of about 64.7-65.4% across efforts, so there is no single max LA cite. Math: ArXivMath 86.8% without tools / 95.2% with tools. Multilingual: GMMLU 92.1% (42 languages), MILU 91.6% (11 languages). Bio capability (not refusal): BioMysteryBench Human Solvable 89.2% / Human Difficult 44.7%; LatchBio SpatialBench Verified 72.5%, SingleCellBench 59.1%; plus System Card §8.17 life-sciences rows (morphology-to-molecule 25.0%, medicinal chemistry ADME 65.3%, protein design sequence 51.0% / library ranking 54.8%, de novo binders 82.3%, biomedical imaging 72.2%, Protocols troubleshooting 67.3% / understanding V2 66.6%).
How to read the table
Interpret by job, not by a single rank. Everyday coding agents: Terminal-Bench, FrontierCode (prefer Xhigh), CursorBench, plus SWE / DeepSWE card depth. Computer use: always pair OSWorld 2.1 partial with strict. Knowledge work: GDPval-AA and AA-Briefcase Elo plus office and AutomationBench rows. Visual / CAD: Chartography and BenchCAD with and without tools. Multidisciplinary reasoning: HLE with and without tools. And keep Anthropic's own split: Sonnet for scoped speed and cost, Opus for hard open-ended judgment.
Pricing & Efficiency
First-party Anthropic list, USD per million tokens. Input and output match Sonnet 5 at $2 / $10, half of Opus 5.5's $4 / $20. Cache reads match Opus at $0.20. The efficiency story is not a sticker cut versus Sonnet 5. It is fewer tokens and faster generation at the same list.
Pricing · vs Opus 5.5
first-party Anthropic list
List price matches Sonnet 5 at $2 / $10, half Opus 5.5's $4 / $20 sticker on input and output. Anthropic says Sonnet 5.5 typically uses fewer tokens and can cost up to 30% less per task than Sonnet 5, while generating outputs 30%+ faster. Batch is 50% off input and output.
Batch discounts are 50% on input and output. Public callers use claude-sonnet-5-5 only.
Effort / Thinking
Adaptive thinking is on by default. Default effort depends on the surface: high on Claude Platform and the API, medium in Claude apps and Claude Code. Effort levels run low → medium → high → xhigh → max. If you migrate from Sonnet 5, re-measure tokens-per-task at the default for your surface before you assume you need max.
Effort & thinking
high on Platform · medium in apps
Adaptive thinking is on by default. Depth is steered with effort, not a separate thinking toggle.
On Claude Platform and the API, default effort is high. That is the everyday production default.
In Claude apps and Claude Code, default effort is medium. Re-measure tokens-per-task if you migrate.
Effort ladder: low, medium, high, xhigh, max. FrontierCode primary in this post is Xhigh 52.1%.
Treat announcement Terminal-Bench carefully: Sonnet's 70.6% sits above Opus 5.5's 66.4% (Opus at Xhigh), but settings may not match. At matched Xhigh, Anthropic's cost framing still places Sonnet as the cheaper lane, not the stronger open-ended model.
When to Use / Migrate
- Good fit: Everyday coding, bug fixes, scoped agent loops, and polished docs / slides / spreadsheets where Sonnet 5 was already the default and you want a clear upgrade without Opus pricing.
- Migrate from Sonnet 5when coding and knowledge-work quality dominate. Same list price; expect fewer tokens and faster outputs on Anthropic's account. Watch default effort (high on Platform, medium in apps).
- Stay on / prefer Opus 5.5for complex open-ended judgment, long-horizon ambiguous agent work, and cases where Anthropic's own guidance says Opus remains clearly stronger. Terminal-Bench alone is not a reason to demote Opus.
- Cyber / biology: first Sonnet with Opus-class cyber safeguards; biology safeguards remain at the Sonnet 5 level. This post does not reprint safety scoreboards.
Caveats
- All launch benches above are self-reported by Anthropic, not LLM Stats verified.
- Terminal-Bench 4.0 Sonnet 70.6% vs Opus 66.4% may mix Max and Xhigh effort. Do not read it as overall Sonnet superiority. Anthropic says Opus remains clearly stronger at complex open-ended work.
- FrontierCode Main primary is Xhigh 52.1%. Max is 46.2% because code-review subagents timed out or made out-of-scope edits.
- OSWorld 2.1 partial (80.1%, announcement) and strict (43.5%, system card) are the same suite; do not cite partial alone. Suite is 2.1, not 2.0.
- Prefer Legal all-pass 11.7% over mean criterion-pass 92.1%. Prefer HealthBench Professional length-adjusted 69.2% (Table 8.1.A; raw 77.1%); PhysicianBench is 63.2%. Plain HealthBench is raw 69.4% with an LA band only. Bio §8.17 rows are capability/knowledge only.
- System Card §8 software-engineering and domain rows are absolute self-reports with no peer deltas in this post. Safety, refusal, and jailbreak tables are intentionally omitted.
- A third-party Elo run for GDPval-AA / AA-Briefcase used a pre-release Platform deployment with a later-fixed structured-outputs bug that may slightly understate scores.
- Knowledge cutoff is June 2026.
Outlook
Anthropic says Claude Haiku 5.5 will follow in the coming weeks. This Sonnet post ages when LLM Stats verifies Terminal-Bench and CursorBench, and when Haiku 5.5 lands.
Through-line for now: second Claude 5.5 model, faster everyday coding and polished knowledge work beside Opus, same Sonnet sticker with fewer tokens and a 30%+ speed claim. For the primary sources, see Anthropic's Sonnet 5.5 announcement, the platform overview, and the system card.
Questions
Frequently Asked Questions
- Anthropic released Claude Sonnet 5.5 on September 28, 2026. It is the second model in the Claude 5.5 family after Opus 5.5, available on the Claude API as
claude-sonnet-5-5. - List price matches Sonnet 5 at $2 / $10 per 1M input / output tokens. Cache reads are $0.20; cache writes are $2.50 (5m) or $4 (1h). Batch is 50% off input and output. Anthropic says fewer tokens can mean up to 30% less cost per task versus Sonnet 5.
- Sonnet 5.5 supports a 1 million token context window with up to 128K max output. Modalities are text + image in, text out. Knowledge cutoff is June 2026.
- Sonnet 5.5 is the everyday complement to Opus 5.5: strong on well-scoped coding and polished knowledge work at half Opus's $4 / $20 sticker. Versus Sonnet 5 it is a clear upgrade on coding and knowledge-work rows, with Anthropic claiming 30%+ faster outputs. Do not treat Terminal-Bench 70.6% vs Opus 66.4% as overall superiority; effort may differ, and Anthropic says Opus remains clearly stronger at complex open-ended work.
- Adaptive thinking is on by default. Default effort is high on Claude Platform / API and medium in Claude apps / Claude Code. Effort levels are low, medium, high, xhigh, and max.
- A self-reported announcement table covering Terminal-Bench 4.0, FrontierCode 1.1 Main, CursorBench 4.0, GDPval-AA v2.1, AA-Briefcase v1.1, HLE with tools, OSWorld 2.1 (partial), and Chartography without tools. The System Card adds secondary capability rows such as SWE-bench Pro 81.3%, DeepSWE v1.1 71.0%, and OSWorld 2.1 strict 43.5% beside the partial 80.1%. FrontierCode primary is 52.1% at Xhigh (Max 46.2%). Scores are vendor-reported, not LLM Stats verified.
Continue Reading
