Claude Opus 5.5: Fable-Class Work, Cheaper to Run
Anthropic's Claude Opus 5.5: Fable 5.1-level on most work, ~40% cheaper than Opus 5 at default. Self-reported Terminal-Bench 66.4%, CursorBench 57.8%, SWE-bench Pro 89.9%; OSWorld strict 48.7% vs partial 81.8%.

Key Numbers
Opus 5.5 · Sep 22, 2026
Efficiency & ID
Fable-class · cheaper to run
Anthropic released Claude Opus 5.5 on September 22, 2026 as the first model in the Claude 5.5 family. The positioning is blunt: perform at Claude Fable 5.1 level on most work, and cost about 40% less to run than Opus 5 at default settings.
That is not a raw score blowout thesis. Anthropic itself says that at these capability levels, benchmark margins are a weaker guide to real-world gaps, and that in their own use the gap to Fable 5.1 is narrower than the table suggests. The buyer story is long-horizon agentic coding efficiency, clearer communication, and a cheaper Opus sticker with a steep cache-read cut.
Headline self-reported numbers from the announcement table: 66.4% on Terminal-Bench 4.0 (xhigh), 57.8% on CursorBench 4.0, 54.4% on FrontierCode v1.1 Main, 67.7% on HLE with tools. List price is $4 / $20 with $0.20 cache reads. API id: claude-opus-5-5. None of the launch scores are LLM Stats verified yet.
At a Glance
- Catalog / API id:
claude-opus-5-5only - Organization: Anthropic
- Release: September 22, 2026. First Claude 5.5 family model.
- Pricing: $4 input / $0.20 cache read / $5 (5m) or $8 (1h) cache write / $20 output per 1M. Batch 50% off.
- Context: 1M input / 128K max output (Batch beta 300K with header)
- Modalities: Text + image in, text out
- Thinking: Adaptive, always on. Default effort
medium - Knowledge cutoff: June 2026
- Safeguards: Production bio / cyber class similar to Fable 5.1
- Secondary card depth: SWE-bench Pro 89.9%, DeepSWE v1.1 74.2%; OSWorld 2.0 strict 48.7% vs partial 81.8%
What's New
Opus 5.5 is Anthropic's first release since the company argued for pacing the frontier. External evaluators named on the launch page include Frontier Design and METR. Anthropic says Opus 5.5 is its strongest model to date on the automated behavioral audit, and that it is less likely than recent models to take hard-to-reverse actions or step outside given boundaries.
Capability claims that matter for buyers: stronger long-horizon coding (migrations, audits, overnight agent runs), clearer writing that puts the important information first, and a cost stack that undercuts Opus 5 on both sticker and tokens-per-task. Anthropic also reports more than 30% faster output generation than Opus 5.
Safety deployment is Fable-class, not Mythos-open. Because Opus 5.5 is comparable to Mythos 5.1 in biology and cybersecurity on Anthropic's account, it ships with safeguards similar to Fable 5.1. Cyber work often falls back to Opus 4.8; biology and frontier LLM development tasks can fall back to Opus 5. Vetted orgs can apply to Life Sciences Verification; Cyber Verification expands in the coming weeks.
Platform docs call out four breaking changes versus Opus 5 code paths: thinking cannot be disabled, forced tool use returns an error, thinking blocks are tied to the model and conversation, and the older computer_20251124 computer-use tool is not accepted on the Claude API and Google Cloud. The first three also apply on Fable 5.1.
Benchmarks
Prefer Anthropic's announcement table for the hero read. Scores below are self-reported, with production safeguards on unless noted. Adaptive thinking at max unless noted; Terminal-Bench 4.0 is the xhigh exception. Not LLM Stats verified.
Opus 5.5 vs priors · percent
safeguards on · SE ±2.6
Zapier · no fallback
SE ±3.5–5
Coding and agents
Terminal-Bench 4.0 at 66.4%(xhigh) is the clearest agentic coding lift versus Fable 5.1 (55.8%) and Opus 5 (52.3%). Standard error is about ±2.6 points for Opus 5.5. Safeguards were on: when they intervened, cybersecurity tasks completed under Opus 4.8, and biology / frontier LLM development tasks under Opus 5. Anthropic says that likely reduces Opus 5.5's score on this harness.
FrontierCode v1.1 Main at 54.4% and CursorBench 4.0 at 57.8%continue the coding story. The announcement body adds effort nuance you should not ignore: at default medium, CursorBench is 52.5% versus Fable 5.1 max 51.8% and Opus 5 max 46.6%. FrontierCode at medium is about 54.6% in the body, versus the table's Main 54.4% at max. Medium is where Anthropic wants the cost story to land.
The System Card adds software-engineering depth behind that agent table: SWE-bench Pro 89.9%, Multilingual 93.9%, Multimodal 61.4%, DeepSWE v1.1 74.2%, FrontierSWE V2 62.3% (Proximal harness, mean across trials), FrontierCode Extended 63.6%, and ProgramBench 91.2%. Card notes best-at-medium FrontierCode Main 54.6% / Extended 65.3%. Those are absolute card rows, not announcement vs-prior deltas.
Knowledge work and computer use
Knowledge work · Elo
GDPval-AA v2.1
announcement table · knowledge work
AA-Briefcase v1.1
system card §8 · bare title
Fable 5.1 · GDPval-AA
same announcement table
GDPval-AA v2.1 at 1846 Elo is the announcement knowledge-work row (bare title only). Fable 5.1 sits at 1735 and Opus 5 at 1708 on the same table. The System Card adds AA-Briefcase v1.1 at 1822 Elo (bare title only). Elo is not a percent bench, so it stays out of the launch bar chart.
AutomationBench at 40.0%(Zapier) improves on Fable 5.1's 31.4% and Opus 5's 26.9%. Important harness caveat: these runs used no fallback models, so safeguard interventions counted as failures. Anthropic says that understates what Opus 5.5 would score in practice.
HLE with tools at 67.7%is a modest hop over Fable 5.1's 65.6% and Opus 5's 63.6%. The card also reports HLE without tools at 64.4%. Terminal-Bench-Science 0.1 at 58.7% keeps the science agent row ahead of Fable 5.1 (52.6%) and well ahead of Opus 5 (29.0%), with SE about ±3.5 to 5 points.
Computer use needs both OSWorld numbers. The announcement partial score is 81.8%; the System Card strict score is 48.7% on the same suite. That gap is the honest read. Chartography with tools is 89.0% on the launch table; without tools the card is 64.4%.
System card domains
Below are absolute System Card §8 capability rows, grouped by job. Safety, refusal, and jailbreak rows are skipped. No invented peer deltas. Self-reported; not LLM Stats verified.
System card §8 · by job
absolute · no peer rows
Software engineering
Depth behind the agentic coding table. Main FrontierCode stays 54.4%.
Proximal harness · mean across trials
Main 54.4%; card best-at-medium Main 54.6% / Ext 65.3%
Computer use
Same OSWorld 2.0 suite as the announcement partial row.
partial on launch table is 81.8%
Vision / CAD
Tool access moves Chartography and BenchCAD a lot.
with tools 89.0% on launch table
Knowledge / office / tools
Office and tool-use depth. Legal all-pass is the hard metric.
all-pass; card mean criterion-pass 91.2%
Health
Prefer length-adjusted over raw when you cite one number.
length-adjusted; raw 77.1%
length-adjusted; raw 68.1%
Math
ArXivMath with and without tools.
Multilingual
Broad MMLU-style coverage across languages.
42 languages
11 languages
Bio capability
Capability and knowledge only. Not refusal or jailbreak rows.
Vision and CAD: BenchCAD is 73.0% without tools and 96.2% with a Python tool. Knowledge and office: OfficeQA 78.9% / OfficeQA Pro 67.7%; Legal Agent Benchmark all-pass 8.3% (card mean criterion-pass 91.2%; prefer all-pass as the standard metric); Toolathlon Verified Pass@1 77.8%. Health: prefer length-adjusted HealthBench Professional 65.6% (raw 77.1%) and HealthBench 60.6% (raw 68.1%). Math: ArXivMath 91.2% without tools / 96.9% with tools. Multilingual: GMMLU 94.3% (42 languages), MILU 93.1% (11 languages). Bio capability (not refusal): BioMysteryBench Human Solvable 89.3% / Human Difficult 50.0%; LatchBio SpatialBench Verified 72.0%, SingleCellBench 61.2%.
How to read the table
Interpret by job, not by a single rank. Coding agents: Terminal-Bench, FrontierCode, CursorBench, plus SWE / DeepSWE card depth. Computer use: always pair OSWorld partial with strict. Knowledge work: GDPval-AA and AA-Briefcase Elo plus AutomationBench and office rows. Science: Terminal-Bench-Science. Visual / CAD: Chartography and BenchCAD with and without tools. Multidisciplinary reasoning: HLE with and without tools. Anthropic's own caveat still applies: margins at this level can overstate day-to-day product gaps versus Fable 5.1.
Pricing & Fast Mode
First-party Anthropic list, USD per million tokens. Input and output are 20% below Opus 5. Cache reads are 60% below Opus 5, which is the line that usually dominates long agent sessions. Anthropic's indexed claim nets the sticker cut plus fewer tokens per task into about 40% lower cost than Opus 5 at default settings.
Pricing · vs Opus 5
first-party Anthropic list
Anthropic's indexed claim: at default settings, Opus 5.5 costs about 40% less than Opus 5 on typical workloads, from cheaper tokens and fewer tokens per step. Batch is 50% off input and output. Fast mode (Claude Code / Platform preview) is $8 / $40 up to 2.5× speed; there is no separate public API model id.
Batch discounts are 50% on input and output. Platform docs list the same 1M context at standard rates with no long-context surcharge called out for this SKU. Public callers use claude-opus-5-5 only.
Effort / Thinking
Adaptive thinking is always on. You do not get a thinking-off path. Default effort on the Claude API is medium. The announcement percent table is mostly max; Terminal-Bench 4.0 is reported at xhigh. If you are migrating from Opus 5 or Fable 5.1, re-measure token-per-task at medium before you assume you need max.
Effort & thinking
adaptive · default medium
Adaptive thinking cannot be disabled. Depth is steered with effort, not a thinking toggle.
API default effort is medium. Several launch charts show medium beating prior max scores at lower cost.
Announcement percent rows are max unless noted. Terminal-Bench 4.0 is the xhigh exception.
Cache reads are 60% below Opus 5. That line dominates agentic and coding session cost.
At default medium, Anthropic reports CursorBench 4.0 at 52.5% versus Fable 5.1 max 51.8% and Opus 5 max 46.6%. FrontierCode at medium is about 54.6% in the body (table Main is 54.4% at max).
When to Use / Migrate
- Good fit:Long-horizon agentic coding, overnight repo work, cache-heavy Claude Code loops, and teams that wanted Fable-class quality without Fable's $10 / $50 sticker.
- Migrate from Opus 5 when cost and tokens-per-task dominate. Expect breaking changes around always-on thinking, forced tool use, preserved thinking blocks, and the older computer-use tool id on some hosts.
- Stay on / prefer Fable 5.1 only if you already optimized around Fable defaults (high effort in Claude Code) and your evals do not show a clear Opus 5.5 win at medium. Anthropic says the real gap can be narrower than the table.
- Cyber / research biology: production safeguards are Fable-class. Expect fallbacks (cyber toward Opus 4.8; bio / frontier LLM development toward Opus 5) unless you are in a verification program.
Caveats
- All launch benches above are self-reported by Anthropic, not LLM Stats verified.
- Terminal-Bench 4.0 used production safeguards; cyber may fall back to Opus 4.8 and bio / frontier LLM development to Opus 5, which likely reduces the headline score.
- AutomationBench (Zapier) ran with no fallback; safeguard interventions counted as failures.
- Terminal-Bench-Science SE is about ±3.5 to 5 points; treat 58.7 as a band.
- OSWorld 2.0 partial (81.8%, announcement) and strict (48.7%, system card) are the same suite; do not cite partial alone.
- System Card §8 software-engineering and domain rows are absolute self-reports with no peer deltas in this post.
- Adaptive thinking cannot be turned off. Knowledge cutoff is June 2026.
- The system card covers safety color and deployment decisions. It is not a source of public numeric safety leaderboard scores in this post.
Outlook
Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks with many of the same performance, efficiency, and safety improvements. This Opus post ages when LLM Stats verifies Terminal-Bench and CursorBench, when mid-tier Claude 5.5 SKUs land, and when preview surfaces (including latency-oriented previews) stabilize in docs.
Through-line for now: first Claude 5.5 model, Fable-class work at a cheaper Opus run cost, medium-default efficiency as the real product claim. For the primary sources, see Anthropic's Opus 5.5 announcement, the platform overview, and the system card.
Questions
Frequently Asked Questions
- Anthropic released Claude Opus 5.5 on September 22, 2026. It is available on the Claude API as
claude-opus-5-5, plus Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. - List price is $4 / $20 per 1M input / output tokens. Cache reads are $0.20; cache writes are $5 (5m) or $8 (1h). Batch is 50% off input and output. Anthropic estimates about 40% lower cost than Opus 5 on typical workloads at default settings.
- Opus 5.5 supports a 1 million token context window with up to 128K max output on the Messages API. Message Batches can go to 300K output with the
output-300k-2026-03-24beta header. Modalities are text + image in, text out. Knowledge cutoff is June 2026. - Anthropic positions Opus 5.5 at Fable 5.1 level on most work, with a clearer efficiency story versus Opus 5. Self-reported Terminal-Bench 4.0 is 66.4%vs Fable 5.1's 55.8% and Opus 5's 52.3%. CursorBench 4.0 is 57.8% vs 51.8% / 46.6%. The buyer case is long-horizon agentic coding efficiency and clearer communication, not a raw score blowout versus Fable 5.1.
- Yes. Adaptive thinking is always on and cannot be disabled. Default effort on the Claude API is medium. Control depth with the effort parameter. Several announcement-body charts highlight medium effort results alongside the max-effort table.
- No. Fast mode is available as a Claude Code / Claude Platform preview at $8 / $40 per 1M with up to 2.5× speed. The public API model id remains
claude-opus-5-5only. - A self-reported announcement table covering Terminal-Bench 4.0, FrontierCode v1.1 Main, CursorBench 4.0, GDPval-AA v2.1, AutomationBench, HLE with tools, Terminal-Bench-Science 0.1, OSWorld 2.0 (partial), and Chartography with tools. The System Card adds secondary capability rows such as SWE-bench Pro 89.9%, DeepSWE v1.1 74.2%, and OSWorld 2.0 strict 48.7% beside the partial 81.8%. Scores are vendor-reported, not LLM Stats verified.
Continue Reading
