The AI arena is free today

Open Superagent
Back to blog
Model Release·Technical Deep Dive

Claude Sonnet 5.5: Faster Everyday Coding Beside Opus 5.5

Anthropic's Claude Sonnet 5.5: faster, cheaper complement to Opus 5.5. Self-reported Terminal-Bench 70.6%, CursorBench 55.5%, FrontierCode Xhigh 52.1%; SWE-bench Pro 81.3%; OSWorld 2.1 strict 43.5% vs partial 80.1%. Same $2/$10 list as Sonnet 5.

Sebastian Crossa
Sebastian Crossa
Co-Founder @ LLM Stats
·10 min read
Claude Sonnet 5.5: Faster Everyday Coding Beside Opus 5.5

Key Numbers

Sonnet 5.5 · Sep 28, 2026

0.0%
Terminal-Bench 4.0
0.0%
CursorBench 4.0
0.0%
FrontierCode Main (Xhigh)
0
GDPval-AA v2.1
$0 / $10
List price / 1M
0%+
Faster than Sonnet 5

Efficiency & ID

faster · cheaper complement

up to 30%
less cost per task vs Sonnet 5
$0.20
Cache reads / MTok
claude-sonnet-5-5
API id
Self-reported announcement table. FrontierCode Main primary is Xhigh 52.1% (Max 46.2% is lower due to code-review subagent timeouts / out-of-scope edits). Speed and cost-per-task claims are Anthropic's vs Sonnet 5. Not LLM Stats verified.

Anthropic released Claude Sonnet 5.5 on September 28, 2026 as the second model in the Claude 5.5 family, after Opus 5.5. The positioning is a complement, not a crown: faster and cheaper for well-scoped everyday work, while Opus keeps the complex open-ended lane.

That buyer story matters more than any single bench row. Sonnet 5.5 is strongest on everyday coding, bug fixes, and polished docs, slides, and spreadsheets (Anthropic highlights a strong design eye). It is a clear upgrade over Sonnet 5 on coding and knowledge work, with the same $2 / $10 list price and a claim of fewer tokens per task.

Headline self-reported numbers from the announcement table: 70.6% on Terminal-Bench 4.0, 55.5% on CursorBench 4.0, 52.1% on FrontierCode 1.1 Main at Xhigh, 64.5% on HLE with tools. List price is $2 / $10 with $0.20 cache reads. API id: claude-sonnet-5-5. None of the launch scores are LLM Stats verified yet.


At a Glance

  • Catalog / API id: claude-sonnet-5-5 only
  • Organization: Anthropic
  • Release: September 28, 2026. Second Claude 5.5 family model after Opus 5.5. Haiku 5.5 still ahead.
  • Pricing: $2 input / $0.20 cache read / $2.50 (5m) or $4 (1h) cache write / $10 output per 1M. Batch 50% off. Same list as Sonnet 5.
  • Context: 1M input / 128K max output
  • Modalities: Text + image in, text out
  • Thinking: Adaptive, on by default. Default effort high on Platform / API; medium in Claude apps / Claude Code
  • Knowledge cutoff: June 2026
  • Safeguards: First Sonnet with cyber safeguards comparable to Opus-class; biology safeguards same as Sonnet 5
  • Secondary card depth: SWE-bench Pro 81.3%, DeepSWE v1.1 71.0%; OSWorld 2.1 strict 43.5% vs partial 80.1%

What's New

Sonnet 5.5 is Anthropic's mid-tier follow-through after Opus 5.5. The product claim is speed and cost for the work most teams actually run every day: scoped coding tasks, bug fixes, and polished office artifacts. Anthropic reports outputs generate 30%+ faster than Sonnet 5, with typically fewer tokens and up to 30% less cost per task at the same list price.

Capability-wise, the announcement table shows large lifts over Sonnet 5 on agentic coding and knowledge-work Elo, and CursorBench within about two points of Opus 5.5. That is the "close enough for most jobs" story, not a claim that Sonnet replaces Opus.

On safeguards, Anthropic says this is the first Sonnet with cyber safeguards comparable to Opus-class models. Cyber capabilities jumped; biology safeguards stay at the Sonnet 5 level. This post does not reprint safety, refusal, or jailbreak tables.


Benchmarks

Prefer Anthropic's announcement table for the hero read. Scores below are self-reported. Not LLM Stats verified. Interpret by job, and keep the Terminal-Bench honesty flag in view before you rank Sonnet above Opus.

Sonnet 5.5 vs peers · percent

Sonnet 5.5Sonnet 5
Terminal-Bench 4.0see note
70.610.3+60.3

Opus at Xhigh · Sonnet may be Max · do not treat as matched effort

FrontierCode 1.1 MainXhigh
52.142.4+9.7

primary Xhigh; Max is 46.2% (timeouts / out-of-scope)

CursorBench 4.0announcement
55.534.1+21.4
HLE (with tools)announcement
64.554.9+9.6
OSWorld 2.1 (partial)announcement
80.157.0+23.1

suite 2.1 · not 2.0

Chartography (no tools)announcement
61.615.6+46.0
Self-reported by Anthropic. Delta column is Sonnet 5.5 minus Sonnet 5. Scores on a 0-100 scale. GDPval-AA and AA-Briefcase Elo are omitted (not percent benches). Terminal-Bench Sonnet vs Opus may mix Max and Xhigh; Anthropic still says Opus is clearly stronger at complex open-ended work. Not LLM Stats verified.

Coding and agents

Terminal-Bench 4.0 at 70.6%is the loudest coding headline versus Sonnet 5's 10.3% and Opus 5.5's 66.4% (Opus at Xhigh). Treat the Sonnet-vs-Opus margin carefully: third-party and Anthropic notes flag a possible Max-versus-Xhigh mismatch. At matched Xhigh, Anthropic's cost framing still places Sonnet lower. Do not claim Sonnet is overall stronger than Opus. Anthropic says Opus remains clearly stronger at complex open-ended work.

FrontierCode 1.1 Main at 52.1% (Xhigh) is the primary row in this post, versus Sonnet 5 at 42.4% and Opus 5.5 at 54.4%. Max on Sonnet is 46.2%, lower than Xhigh because code-review subagents hit timeouts and out-of-scope edits. CursorBench 4.0 at 55.5%lands within about two points of Opus 5.5's 57.8%, and well above Sonnet 5's 34.1%.

The System Card adds software-engineering depth behind that agent table: SWE-bench Pro 81.3%, Multilingual 90.3%, Multimodal 54.3%, DeepSWE v1.1 71.0%, FrontierSWE V2 61.9% (Proximal harness, mean across trials), FrontierCode Extended 64.4% at Xhigh, and ProgramBench 79.7%. Terminal-Bench-Science 0.1 is 59.9% on the card. Those are absolute card rows, not announcement vs-prior deltas.

Knowledge work and computer use

Knowledge work · Elo

1844

GDPval-AA v2.1

Sonnet 5.5 · announcement table

1811

AA-Briefcase v1.1

Sonnet 5.5 · announcement table

1846

Opus 5.5 · GDPval-AA

same announcement table · within 2 Elo

Sonnet 5 sits at 1449 GDPval-AA and 1359 AA-Briefcase on the same table. Opus 5.5 is 1846 / 1822. Sonnet 5.5 lands within a few Elo of Opus on both rows while keeping the Sonnet sticker.

Elo, not percent. Bare titles only from Anthropic's announcement table. A third-party Elo run used a pre-release Platform deployment with a later-fixed structured-outputs bug that may slightly understate scores. Self-reported. Not LLM Stats verified.

GDPval-AA v2.1 at 1844 Elo and AA-Briefcase v1.1 at 1811 Elo are the announcement knowledge-work rows (bare titles only). Both sit within a few Elo of Opus 5.5 (1846 / 1822) and far above Sonnet 5 (1449 / 1359). Elo is not a percent bench, so it stays out of the launch bar chart.

HLE with tools at 64.5%improves on Sonnet 5's 54.9% and trails Opus 5.5's 67.7%. The card also reports HLE without tools at 56.9%. Computer use needs both OSWorld 2.1 numbers: the announcement partial score is 80.1%; the System Card strict score is 43.5% on the same suite. That gap is the honest read. Chartography without tools is 61.6% on the launch table; with tools the card is 90.2%.

System card domains

Below are absolute System Card §8 capability rows, grouped by job. Safety, refusal, and jailbreak rows are skipped. No invented peer deltas. Self-reported; not LLM Stats verified.

System card §8 · by job

absolute · no peer rows

Software engineering

Depth behind the agentic coding table. Main FrontierCode stays 52.1% Xhigh.

SWE-bench Pro81.3%
SWE-bench Multilingual90.3%
SWE-bench Multimodal54.3%
DeepSWE v1.171.0%
FrontierSWE V261.9%

Proximal harness · mean across trials

FrontierCode v1.1 Extended64.4%

Xhigh; Main Xhigh 52.1% / Max Main 46.2%

ProgramBench79.7%

Computer use

Same OSWorld 2.1 suite as the announcement partial row.

OSWorld 2.1 strict43.5%

partial on launch table is 80.1%

Vision / CAD

Tool access moves Chartography and BenchCAD a lot.

Chartography (no tools)61.6%

with tools 90.2%

BenchCAD (no tools)74.7%

voxel IoU

BenchCAD (Python tool)96.3%

Knowledge / office / tools

Office and tool-use depth. Legal all-pass is the hard metric.

OfficeQA76.9%
OfficeQA Pro65.6%
Legal Agent Benchmark11.7%

all-pass; card mean criterion-pass 92.1%

Toolathlon Verified Pass@177.8%

Health

Prefer Table 8.1.A length-adjusted Professional when citing one number.

HealthBench Professional69.2%

length-adjusted Table 8.1.A; raw 77.1%

PhysicianBench63.2%

100 physician EHR tasks · max effort

HealthBench (raw)69.4%

LA across efforts is a 64.7-65.4% band only

Math

ArXivMath with and without tools.

ArXivMath (no tools)86.8%
ArXivMath (with tools)95.2%

Multilingual

Broad MMLU-style coverage across languages.

GMMLU92.1%

42 languages

MILU91.6%

11 languages

Bio capability

Capability and knowledge only, including System Card §8.17. Not refusal or jailbreak rows.

BioMysteryBench Human Solvable89.2%
BioMysteryBench Human Difficult44.7%
LatchBio SpatialBench Verified72.5%
LatchBio SingleCellBench59.1%
Morphology-to-molecule matching25.0%

Axiom Bio

Medicinal Chemistry ADME65.3%

504 questions

Protein Design Sequence Generation51.0%
Protein Design Library Ranking54.8%
De novo protein binder design82.3%
Biomedical image analysis72.2%
Protocols Troubleshooting67.3%
Protocols Understanding (V2)66.6%

Benchling

Self-reported Anthropic System Card §8 capability rows for Sonnet 5.5 only. Absolute bars, not vs-prior deltas. Safety, refusal, and jailbreak rows are omitted. Elo benches are in EloStory. Not LLM Stats verified.

Vision and CAD: BenchCAD is 74.7% without tools and 96.3% with a Python tool. Knowledge and office: OfficeQA 76.9% / OfficeQA Pro 65.6%; Legal Agent Benchmark all-pass 11.7% (card mean criterion-pass 92.1%; prefer all-pass as the standard metric); Toolathlon Verified Pass@1 77.8%. AutomationBench is 44.7% (Zapier, max effort, API default fallbacks enabled). Health: prefer length-adjusted HealthBench Professional 69.2% from Table 8.1.A (raw 77.1%); PhysicianBench is 63.2% (100 physician EHR tasks, max effort); HealthBench is stored as raw 69.4% with an LA band of about 64.7-65.4% across efforts, so there is no single max LA cite. Math: ArXivMath 86.8% without tools / 95.2% with tools. Multilingual: GMMLU 92.1% (42 languages), MILU 91.6% (11 languages). Bio capability (not refusal): BioMysteryBench Human Solvable 89.2% / Human Difficult 44.7%; LatchBio SpatialBench Verified 72.5%, SingleCellBench 59.1%; plus System Card §8.17 life-sciences rows (morphology-to-molecule 25.0%, medicinal chemistry ADME 65.3%, protein design sequence 51.0% / library ranking 54.8%, de novo binders 82.3%, biomedical imaging 72.2%, Protocols troubleshooting 67.3% / understanding V2 66.6%).

How to read the table

Interpret by job, not by a single rank. Everyday coding agents: Terminal-Bench, FrontierCode (prefer Xhigh), CursorBench, plus SWE / DeepSWE card depth. Computer use: always pair OSWorld 2.1 partial with strict. Knowledge work: GDPval-AA and AA-Briefcase Elo plus office and AutomationBench rows. Visual / CAD: Chartography and BenchCAD with and without tools. Multidisciplinary reasoning: HLE with and without tools. And keep Anthropic's own split: Sonnet for scoped speed and cost, Opus for hard open-ended judgment.


Pricing & Efficiency

First-party Anthropic list, USD per million tokens. Input and output match Sonnet 5 at $2 / $10, half of Opus 5.5's $4 / $20. Cache reads match Opus at $0.20. The efficiency story is not a sticker cut versus Sonnet 5. It is fewer tokens and faster generation at the same list.

Pricing · vs Opus 5.5

first-party Anthropic list

$2
Opus $4
Input
per 1M · same as Sonnet 5
$10
Opus $20
Output
per 1M · same as Sonnet 5
$0.20
Opus $0.20
Cache read
per 1M · matched Opus line
$2.50 / $4
Opus $5 / $8
Cache write
5m / 1h vs Opus 5.5

List price matches Sonnet 5 at $2 / $10, half Opus 5.5's $4 / $20 sticker on input and output. Anthropic says Sonnet 5.5 typically uses fewer tokens and can cost up to 30% less per task than Sonnet 5, while generating outputs 30%+ faster. Batch is 50% off input and output.

USD per million tokens. Cache writes: $2.50 (5m) / $4 (1h) on Sonnet 5.5 vs $5 / $8 on Opus 5.5. Public API id remains claude-sonnet-5-5.

Batch discounts are 50% on input and output. Public callers use claude-sonnet-5-5 only.


Effort / Thinking

Adaptive thinking is on by default. Default effort depends on the surface: high on Claude Platform and the API, medium in Claude apps and Claude Code. Effort levels run low → medium → high → xhigh → max. If you migrate from Sonnet 5, re-measure tokens-per-task at the default for your surface before you assume you need max.

Effort & thinking

high on Platform · medium in apps

Adaptive on

Adaptive thinking is on by default. Depth is steered with effort, not a separate thinking toggle.

Default high

On Claude Platform and the API, default effort is high. That is the everyday production default.

Apps medium

In Claude apps and Claude Code, default effort is medium. Re-measure tokens-per-task if you migrate.

low → max

Effort ladder: low, medium, high, xhigh, max. FrontierCode primary in this post is Xhigh 52.1%.

Treat announcement Terminal-Bench carefully: Sonnet's 70.6% sits above Opus 5.5's 66.4% (Opus at Xhigh), but settings may not match. At matched Xhigh, Anthropic's cost framing still places Sonnet as the cheaper lane, not the stronger open-ended model.

From Anthropic platform docs and the Sonnet 5.5 announcement. Conceptual rails only. Default effort differs by surface: high on Platform / API, medium in Claude apps and Claude Code.

When to Use / Migrate

  • Good fit: Everyday coding, bug fixes, scoped agent loops, and polished docs / slides / spreadsheets where Sonnet 5 was already the default and you want a clear upgrade without Opus pricing.
  • Migrate from Sonnet 5when coding and knowledge-work quality dominate. Same list price; expect fewer tokens and faster outputs on Anthropic's account. Watch default effort (high on Platform, medium in apps).
  • Stay on / prefer Opus 5.5for complex open-ended judgment, long-horizon ambiguous agent work, and cases where Anthropic's own guidance says Opus remains clearly stronger. Terminal-Bench alone is not a reason to demote Opus.
  • Cyber / biology: first Sonnet with Opus-class cyber safeguards; biology safeguards remain at the Sonnet 5 level. This post does not reprint safety scoreboards.

Caveats

  • All launch benches above are self-reported by Anthropic, not LLM Stats verified.
  • Terminal-Bench 4.0 Sonnet 70.6% vs Opus 66.4% may mix Max and Xhigh effort. Do not read it as overall Sonnet superiority. Anthropic says Opus remains clearly stronger at complex open-ended work.
  • FrontierCode Main primary is Xhigh 52.1%. Max is 46.2% because code-review subagents timed out or made out-of-scope edits.
  • OSWorld 2.1 partial (80.1%, announcement) and strict (43.5%, system card) are the same suite; do not cite partial alone. Suite is 2.1, not 2.0.
  • Prefer Legal all-pass 11.7% over mean criterion-pass 92.1%. Prefer HealthBench Professional length-adjusted 69.2% (Table 8.1.A; raw 77.1%); PhysicianBench is 63.2%. Plain HealthBench is raw 69.4% with an LA band only. Bio §8.17 rows are capability/knowledge only.
  • System Card §8 software-engineering and domain rows are absolute self-reports with no peer deltas in this post. Safety, refusal, and jailbreak tables are intentionally omitted.
  • A third-party Elo run for GDPval-AA / AA-Briefcase used a pre-release Platform deployment with a later-fixed structured-outputs bug that may slightly understate scores.
  • Knowledge cutoff is June 2026.

Outlook

Anthropic says Claude Haiku 5.5 will follow in the coming weeks. This Sonnet post ages when LLM Stats verifies Terminal-Bench and CursorBench, and when Haiku 5.5 lands.

Through-line for now: second Claude 5.5 model, faster everyday coding and polished knowledge work beside Opus, same Sonnet sticker with fewer tokens and a 30%+ speed claim. For the primary sources, see Anthropic's Sonnet 5.5 announcement, the platform overview, and the system card.

Questions

Frequently Asked Questions

  • Anthropic released Claude Sonnet 5.5 on September 28, 2026. It is the second model in the Claude 5.5 family after Opus 5.5, available on the Claude API as claude-sonnet-5-5.
  • List price matches Sonnet 5 at $2 / $10 per 1M input / output tokens. Cache reads are $0.20; cache writes are $2.50 (5m) or $4 (1h). Batch is 50% off input and output. Anthropic says fewer tokens can mean up to 30% less cost per task versus Sonnet 5.
  • Sonnet 5.5 supports a 1 million token context window with up to 128K max output. Modalities are text + image in, text out. Knowledge cutoff is June 2026.
  • Sonnet 5.5 is the everyday complement to Opus 5.5: strong on well-scoped coding and polished knowledge work at half Opus's $4 / $20 sticker. Versus Sonnet 5 it is a clear upgrade on coding and knowledge-work rows, with Anthropic claiming 30%+ faster outputs. Do not treat Terminal-Bench 70.6% vs Opus 66.4% as overall superiority; effort may differ, and Anthropic says Opus remains clearly stronger at complex open-ended work.
  • Adaptive thinking is on by default. Default effort is high on Claude Platform / API and medium in Claude apps / Claude Code. Effort levels are low, medium, high, xhigh, and max.
  • A self-reported announcement table covering Terminal-Bench 4.0, FrontierCode 1.1 Main, CursorBench 4.0, GDPval-AA v2.1, AA-Briefcase v1.1, HLE with tools, OSWorld 2.1 (partial), and Chartography without tools. The System Card adds secondary capability rows such as SWE-bench Pro 81.3%, DeepSWE v1.1 71.0%, and OSWorld 2.1 strict 43.5% beside the partial 80.1%. FrontierCode primary is 52.1% at Xhigh (Max 46.2%). Scores are vendor-reported, not LLM Stats verified.

Continue Reading