Back to blog
Comparison·Benchmarks·2026 Guide

Claude Opus 5 vs Claude Fable 5: Ultimate Comparison

Claude Opus 5 vs Claude Fable 5 on benchmarks, pricing, safeguards, and API tradeoffs. The cheaper Opus wins more comparable launch-table rows.

Jonathan Chavez
Jonathan Chavez
Co-Founder @ LLM Stats
·9 min read
Claude Opus 5 vs Claude Fable 5: Ultimate Comparison

Anthropic's two top general-access Claude tiers are easy to misread. Fable 5 launched as the safer public configuration of Mythos 5, carrying a $10 / $50 per million input / output token price. Six weeks later, Claude Opus 5arrived at $5 / $25, with Anthropic positioning it close to Fable's frontier intelligence at half the cost.

The useful comparison is not a claim that one tier wins every task. Anthropic's own tables show Fable ahead on DeepSWE v1.1 and the Legal Agent Benchmark. But Opus leads the directly comparable published rows for Frontier-Bench, GDPval-AA, BrowseComp, OSWorld 2.0, and AutomationBench. The most important caveat is configuration: several headline figures beside Fable are explicitly Mythos 5 figures, not measurements of the safeguarded public model.


The Verdict

Use Claude Opus 5 as the default.It costs half as much as Fable 5, is the new default on Claude Max, and its public launch table beats Fable on five of nine directly comparable percentage benchmarks. That is not a small detail. It means the lower-priced model is already better on agentic terminal coding, search, computer use, and business workflows in Anthropic's current comparison.

Use Fable 5 deliberately, not by default.Its 69.7% on DeepSWE v1.1 edges Opus's 68.8%, and its 13.3% on the Legal Agent Benchmark held-out set exceeds Opus's 11.7%. If either workload maps closely to production traffic, test both models. Those wins may justify the premium. Otherwise, Opus offers the more convincing capability per dollar.

Do not select Fable from Mythos figures. The Fable 5 launch announcement states that its safeguards can route cybersecurity, biology, chemistry, and distillation queries to Opus 4.8. The correct question is what the public Fable configuration does on your prompts, not what the less-restricted Mythos configuration did in a benchmark run.


At a Glance

DetailClaude Opus 5Claude Fable 5
API model IDclaude-opus-5claude-fable-5
ReleasedJuly 24, 2026June 9, 2026
Input / output price$5 / $25 per 1M$10 / $50 per 1M
Input / output context1M / 128K tokens1M / 128K tokens
ModalitiesText and image in, text outText and image in, text out
PositioningEveryday flagship and Claude Max defaultSafeguarded Mythos-class specialist
Guarded requestsOpus 5 response, subject to its own policyCan fall back to Opus 4.8

How to Read the Benchmark Table

This is not a clean lab shootout. Both scorecards are published by Anthropic, at different times, with benchmark-specific harnesses and effort settings. A raw number is meaningful only within its benchmark row. GDPval-AA is an Elo measure, for example, while OSWorld is a task success rate. Neither should be averaged into a synthetic intelligence score.

The configuration split matters even more. The Opus 5 comparison graphic labels some Fable-side values as Mythos 5, specifically HealthBench Professional and both BioMysteryBench rows. Those figures are useful evidence about the shared underlying model, but they are not evidence that a public Fable 5 request will obtain the same result. This article keeps them out of the direct win count.


Benchmark Head-to-Head

Across nine shared percentage rows that name Fable 5 directly, Opus wins five, Fable wins three, and FrontierCode is effectively tied. Opus also leads GDPval-AA v2 at 1861 versus 1747 Elo. The pattern is notable because the cheaper model's largest margins are not narrow ties: it leads Fable by 9.6 points on Frontier-Bench and 8.6 points on AutomationBench.

Published head-to-head

Seven shared percentage scores

Opus leads the published table
where the two can be compared directly.

Frontier-Bench

Opus
43.3
Fable
33.7

Opus 5

BrowseComp

Opus
90.8
Fable
87.4

Opus 5

OSWorld 2.0

Opus
70.6
Fable
66.1

Opus 5

DeepSWE v1.1

Opus
68.8
Fable
69.7

Fable 5

FrontierCode Main

Opus
53.4
Fable
53.5

Near tie

AutomationBench

Opus
26.0
Fable
17.4

Opus 5

Legal Agent Bench

Opus
11.7
Fable
13.3

Fable 5

Anthropic's June and July 2026 launch comparison tables. Each row uses its published setup, so bar length compares models only within a row, not across benchmarks. GDPval-AA is excluded because it is an Elo score.
BenchmarkOpus 5Fable 5Read
Frontier-Bench v0.143.3%33.7%Opus +9.6
GDPval-AA v21861 Elo1747 EloOpus +114
BrowseComp90.8%87.4%Opus +3.4
HLE, no tools56.3%56.5%Fable +0.2
OSWorld 2.070.6%66.1%Opus +4.5
DeepSWE v1.168.8%69.7%Fable +0.9
FrontierCode v1.1 Main53.4%53.5%Near tie
AutomationBench26.0%17.4%Opus +8.6
Legal Agent Benchmark11.7%13.3%Fable +1.6

The part that actually changes a routing decision

Opus's lead is concentrated in broad agent work: terminal coding, search, computer use, knowledge work, and workflow automation. That is exactly where a general-purpose default will receive most of its volume. Fable's published advantage is narrower, but still real, on DeepSWE and legal-agent work. A team building an autonomous code-maintenance or legal workflow should run an evaluation with its own tasks before replacing Fable.

The HLE and FrontierCode margins are not decision signals. A 0.2 or 0.1 point difference is below the resolution needed to change a production route, especially without identical harnesses. Treat them as ties and make the choice on task completion, retries, latency, and token use.


Safeguards Change the Comparison

Fable 5 is not simply a more expensive model ID. Anthropic describes it as the general-access version of Mythos 5, with safeguards intended to prevent misuse. For some cybersecurity, biology, chemistry, and distillation requests, Fable routes to Claude Opus 4.8. Anthropic said these safeguards triggered in fewer than 5% of sessions on average at launch, while acknowledging that conservative classifiers can catch harmless work.

This means that the operational result can differ from the model-card result. A Fable 5 benchmark score may describe the model where the safeguard did not intervene. A production task near a guarded boundary can receive an older Opus answer instead. Opus 5 does not eliminate policy constraints, but it does remove this specific Fable-to-Opus 4.8 downgrade from the selection logic.

That is why the Mythos values should remain in a separate column, not silently folded into Fable's win record. They measure a valuable capability frontier, but they do not describe the same public API behavior.


Pricing and API Surface

The list-price comparison is exact: Fable is 2x Opus at both input and output. At 1 million input and 1 million output tokens, Opus costs $30 and Fable costs $60. On a workload that emits large reasoning traces or performs several retries, that gap grows quickly.

WorkloadOpus 5 list costFable 5 list cost
1M input tokens$5$10
1M output tokens$25$50
1M input + 1M output$30$60

Token price is not task price. A model that completes a difficult task in fewer turns may offset its higher rate. Anthropic makes that case for both releases. But the published benchmark table gives no universal reason to assume Fable will use fewer tokens than Opus. Measure completion rate, retries, output tokens, and human review time on your own traces before treating the premium as justified.


Which Model Should You Use?

  • Default to Opus 5 for general agents.It is half the list price and leads Fable's directly comparable published scores in the broad categories that usually dominate production traffic.
  • Test Fable 5 for DeepSWE-like code maintenance and legal workflows. These are its two clear published wins. Promote it only if the lift survives your evaluation set.
  • Do not route sensitive scientific or cyber work to Fable expecting Mythos performance.The safeguard fallback changes the effective model. Use Anthropic's trusted-access options where applicable and validate the exact supported configuration.
  • Route by task difficulty rather than brand tier. Use a cheaper model for routine work, Opus for hard general work, and Fable only where measured task-level results pay for its premium.

The more interesting finding is not that Opus 5 is cheaper. It is that Anthropic's own current comparison no longer places its public hierarchy in a straight line. Fable remains a valuable specialist tier, but Opus 5 is the broader, cheaper, and more transparent default. For the underlying sources, read Anthropic's Opus 5 announcement and Fable 5 announcement.

Questions

Frequently Asked Questions

  • For most general-access tasks, Claude Opus 5is the stronger default. It costs half as much and leads Fable 5 on five of nine directly comparable percentage rows in Anthropic's published launch tables. Fable 5 does lead on DeepSWE v1.1 and the Legal Agent Benchmark, but its separate Mythos 5 scores are not public Fable 5 scores.
  • Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens. Claude Fable 5 costs $10 input and $50 output. Fable is exactly 2x the listed token price.
  • Mythos 5 and Fable 5 share underlying weights, but Fable 5 adds general-access safeguards. Anthropic says some cybersecurity, biology, chemistry, and distillation requests fall back to Claude Opus 4.8. A Mythos 5 score therefore does not automatically describe what a public Fable 5 API request will achieve.
  • Both models support a 1 million-token input context window and up to 128K output tokens. Both accept text and image input and produce text output.

Continue Reading