Skip to content

Claude Opus 5 Benchmarks: Reading the Numbers Against Fable 5 and GPT-5.6 Sol

Comparison as of July 25

What the data supportsWhat the data does not establish
The composite scores are close at 61, 60, and 59A one-point gap is not a decisive real-world lead
Opus 5 leads professional-work evaluations by more than 100 EloA private workload will reproduce the same gap
Opus 5 list prices are 50% below Fable 5Measured task cost falls by the same 50%
Both Claude evaluations used fallbackPublic scores isolate either base model's pure performance

Anthropic released Claude Opus 5 on July 24, 2026. The pitch is near-Fable 5 intelligence at half the price, with API rates of $5 per million input tokens and $25 per million output tokens.1

Artificial Analysis published independent results by the following day. They support a more qualified conclusion. Composite intelligence is nearly tied, Opus 5 leads professional agentic work, and the measured task-cost reduction is 26%, not 50%.

Claude Opus 5, Fable 5, and GPT-5.6 Sol compared across four decision axes

The professional-work gap matters more than the one-point composite lead

On Artificial Analysis Intelligence Index v4.1, Opus 5 at max effort scores 61, Fable 5 scores 60, and GPT-5.6 Sol scores 59. The composite combines nine evaluations, including GDPval-AA v2, Terminal-Bench v2.1, and Humanity's Last Exam.23

Independent evaluationOpus 5Fable 5GPT-5.6 Sol
Intelligence Index v4.1616059
GDPval-AA v2 Elo1,8611,7471,736
AA-Briefcase Elo1,720about 1,574no published comparison

Opus 5 leads Fable 5 by 114 Elo and Sol by 125 Elo on GDPval-AA v2. Its AA-Briefcase lead over Fable 5 is 146 Elo.2

The one-point Index margin averages across broad capabilities. GDPval-AA v2 and AA-Briefcase focus more directly on producing professional deliverables with files and tools. The defensible reading is therefore a lead in agentic professional work, not a decisive general-intelligence hierarchy.

One independent result cuts against a simple winner narrative. Artificial Analysis reports that Opus 5 improved seven points over Opus 4.8 on AA-Omniscience accuracy, while its hallucination rate rose 14 points to 50%.2 Higher professional-work Elo does not guarantee safe factual recall.

A 50% list-price cut becomes a 26% task-cost cut

Opus 5 and Fable 5 have a clean two-to-one API price ratio.14

Per million tokensOpus 5Fable 5GPT-5.6 Sol
Input$5$10$5
Output$25$50$30

Artificial Analysis measured average Intelligence Index task cost at $2.03 for Opus 5 and $2.75 for Fable 5. That is about 26% lower, not 50% lower.2

The gap narrows because a bill depends on more than list price. Reasoning output, answer output, input, cache writes, and cache hits all vary by model and effort. Artificial Analysis reports roughly an eightfold range in GDPval-AA v2 output-token use from low to max effort.2

A migration plan therefore cannot assume that replacing Fable 5 halves the bill. It needs total tokens, retries, elapsed time, and cost to the same acceptance criterion.

Effort changes latency and token use along with intelligence

Opus 5 offers low, medium, high, xhigh, and max effort. Artificial Analysis measured an Index range from 51 at low to 61 at max, while total evaluation output rose from 12M to 100M tokens.2

The provider-performance figures were moving even during July 25. Artificial Analysis showed xhigh time to first token in a range from roughly 37 to 40 seconds at different retrieval times. Those rolling measurements should not be copied into a production SLA as fixed constants.

The stable signal is that xhigh and max can delay the first response by tens of seconds. Interactive products should start at low or medium and raise effort only for tasks that fail quality criteria. Fast mode is a separate lever: Anthropic advertises about 2.5 times the standard speed for twice the base price.1

Vendor benchmarks require a separate evidence label

Anthropic reports 43.3% for Opus 5 on Frontier-Bench v0.1. Published chart figures put Fable 5 at 33.7%, GPT-5.6 Sol at 34.4%, and Opus 4.8 at 18.7%.15

On ARC-AGI 3, Opus 5 scores 30.2% against Sol's 7.8%. Fable 5 was not tested, so this result cannot establish an Opus-versus-Fable ranking.5

Anthropic also says Opus 5 comes within 0.5% of Fable 5's peak on CursorBench 3.2, delivers about 1.5 times the next-best Zapier AutomationBench pass rate at equal cost, and exceeds Fable 5's best OSWorld 2.0 result at roughly one-third the cost.1 These are useful adoption signals, but Anthropic or its partners produced or presented them. They do not carry the same evidentiary weight as an independent run.

Fallback is part of both Claude evaluation configurations

Fable 5 is not the only composite configuration in the comparison. Artificial Analysis explicitly says its Opus 5 Intelligence Index run enabled server-side fallback to Opus 4.8.2

Anthropic's Frontier-Bench footnote also says Opus 4.8 handled safety-classifier refusals for both Opus 5 and Fable 5. The internal run used mini-SWE-agent on a GKE backend and averaged five attempts per task.1

The public comparison is therefore among:

  • Opus 5 with Opus 4.8 fallback
  • Fable 5 with Opus 4.8 fallback
  • the evaluated GPT-5.6 Sol configuration

Sharing a fallback target does not prove that its effect cancels out. Different trigger rates and task categories can change each aggregate differently. The published results cannot isolate either Claude base model's fallback-free score.

The safeguards move Opus 5 toward everyday professional use

Opus 5 allows vulnerability discovery in source code while blocking binary-based scanning, penetration testing, and exploit generation. Anthropic estimates that its classifiers will intervene about 85% less often than Fable 5's.16

Flagged requests fall back to Opus 4.8 by default in Claude.ai, Claude Code, and Claude Cowork. The API offers automatic fallback as an opt-in beta; without it, the request is blocked. Cyber Verification Program participants can use a less restricted configuration.1

For biology requests, Opus 5 replaces Opus 4.8 as the fallback from Fable 5. General Opus 5 access does not carry Fable 5's additional data-retention requirement.1

Adoption decisions split along three boundaries

  1. Opus 4.8 users can compare at the same price. The Index rises from 56 to 61, while Anthropic's Frontier-Bench figure rises from 18.7% to 43.3%. Opus 5 is a strong upgrade candidate if it passes existing regression tests.
  2. Fable 5 users should model cost per accepted task. The independent gap is 26%, so long-running success rate and retries belong in the calculation.
  3. Sol users should separate professional-work quality from interactive latency. The Index is 61 versus 59 and GDPval-AA v2 differs by 125 Elo, but high Opus 5 effort can impose a long wait.

Use at least 20 representative private tasks. Record pass rate, cost to pass, completion time, and fallback events to turn the public benchmark into an operating decision.


  1. Anthropic, Introducing Claude Opus 5, July 24, 2026. 

  2. Artificial Analysis, Opus 5: Fable 5 level intelligence at a lower cost per task, July 24, 2026. 

  3. Artificial Analysis, Intelligence Benchmarking methodology, accessed July 25, 2026. 

  4. Anthropic, Models overview, accessed July 25, 2026. 

  5. Decrypt, Claude Opus 5 Outscores Fable 5 on Most Benchmarks—At Half the Price, July 24, 2026. Used as a secondary transcription of values displayed in the launch charts. 

  6. Anthropic, Claude Opus 5 System Card, July 24, 2026.