Claude Opus 5: Frontier Scores at Half of Fable's Cost
The Fourth Model in Seven Weeks
Anthropic released Claude Opus 5 on July 24, 2026. It is the fourth model the company has put into the world since June 9, following Mythos 5, Fable 5, and Sonnet 5 — a cadence that would have been implausible eighteen months ago and now barely registers as unusual.
The framing Anthropic chose is unambiguous. Opus 5 comes close to the frontier intelligence of Claude Fable 5 at half the price, and on coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, it is the new state of the art, though it remains behind Mythos 5 on cybersecurity tasks. It becomes the default model on Claude Max and the strongest model available on Claude Pro.
That is a strange sentence for a company to publish about its own lineup. Fable 5 is Anthropic's most capable generally available model. Opus 5 is the cheaper tier below it. And Anthropic is telling you, in the launch post, that the cheaper tier is now the state of the art on the evaluations most buyers care about.
The interesting question is not whether that claim is true. It is what it says about where the industry's constraint has moved. Six months ago the fight was over peak capability. Today it is over cost per completed task, and Anthropic just released a model whose entire pitch is that axis.
What Actually Shipped
The concrete details, before the analysis:
Pricing: $5 per million input tokens and $25 per million output tokens — identical to Opus 4.8. No introductory discount, no price increase. The generational upgrade is free in per-token terms.
Context: A 1M-token context window, which is both the default and the maximum — there is no smaller context variant. Max output is 128k tokens.
Thinking: On by default. This is a behavior change worth flagging: disabling thinking at xhigh or max effort returns a 400 error.
Effort: The full ladder — low, medium, high, xhigh, and max, with max as the explicit top tier for the deepest reasoning. No beta header required.
Model string: claude-opus-5 on the Claude API.
Fast mode: Opus 5 runs at roughly 2.5 times default speed in Fast mode, priced at twice the base rate on the Claude Platform and drawing on usage credits in Claude Code.
Data retention: Consistent with prior Opus models, Opus 5 has no data retention requirements for general access — a meaningful distinction from Fable 5, which carries a 30-day retention condition.
Availability is immediate across Anthropic's own apps, the Claude Platform, and the three cloud partners.
The Benchmark Claims, Read Carefully
Anthropic leans on a specific set of evaluations, and it is worth separating what each one measures.
Frontier-Bench v0.1. Anthropic says Opus 5 surpasses all other models and more than doubles Opus 4.8's performance at a lower cost per task. Doubling a predecessor's score in under two months is the kind of claim that deserves a footnote, and Anthropic supplies one: the results come from an internal run on the mini-SWE-agent harness with a GKE backend, mean reward over five attempts per task, with Opus 4.8 serving as the fallback on safety-classifier refusals for both Opus 5 and Fable 5. That is an unusually transparent disclosure, and it also means the number is a vendor-run figure on a v0.1 benchmark. Treat it as directional until third parties reproduce it.
CursorBench 3.2. At max effort, Opus 5 lands within 0.5% of Fable 5's peak score at half the cost per task, and achieves better performance at a given cost than every other model at high, xhigh, and max effort. This is the cleanest statement of the launch thesis: not "we beat Fable," but "we get within rounding distance of Fable for half the money."
ARC-AGI 3. Anthropic reports Opus 5 scoring three times as high as the next-best model on this novel-problem evaluation. ARC-style benchmarks have historically been where model families either plateau hard or jump discontinuously, so a 3× margin is the single most eyebrow-raising number in the announcement — and the one most worth watching for independent replication.
Zapier AutomationBench. Opus 5's pass rate is around 1.5× the next-best model at the same cost per task, and even at its lowest effort setting it passes more tasks than any other model. That second clause matters more than the first for anyone running high-volume automation.
OSWorld 2.0. On computer use, Opus 5 outperforms every other model at any given cost, and surpasses Fable 5's best result at just over a third of the cost.
Life sciences. Anthropic reports gains on every internal life sciences evaluation versus Opus 4.8, with the largest on organic chemistry — 10.2 percentage points higher on inferring molecular structures from spectroscopy data, and 7.7 points higher on predicting how protein sequence variation affects function.
The pattern across all of these is consistent, and it is not "highest score." It is "highest score per dollar." Every chart Anthropic published plots performance against cost rather than performance alone. That is a deliberate rhetorical choice and it reflects a real shift in what buyers are asking.
The Artificial Analysis Result
Independent confirmation arrived within hours. Claude Opus 5 debuted at the top of the Artificial Analysis Intelligence Index with a score of 61, ahead of Claude Fable 5 at 60 and pushing GPT-5.6 Sol to third at 59.
One point is inside the noise band of any composite index, and anyone treating a 61-vs-60 gap as a decisive capability ordering is reading it wrong. What is not inside the noise band is the cost differential. Fable 5 lists at $10/$50 per million tokens — double Opus 5 — and when Artificial Analysis benchmarked Fable 5 in June, it cost roughly $6.2K to run the Intelligence Index, the most expensive model they had ever benchmarked, at 1.7× the next-highest model.
So the ranking story is a tie. The economics story is not. A model that scores at parity with the frontier while consuming a fraction of the tokens and billing at half the rate is a materially different product, even if the leaderboard row looks nearly identical.
Context for how crowded that leaderboard has become: Moonshot AI's Kimi K3, released July 16 at 2.8 trillion parameters, scored around 57 — fourth overall — while leading the Frontend Code Arena and posting the strongest open-weight GPQA Diamond result to date at 93.5%. The gap between the best closed model and the best open-weight model on general intelligence is now roughly four points. That is the backdrop against which Anthropic is competing on price.
Cost Per Task Is the Real Metric
The per-token price is the number on the pricing page. It is not the number on your invoice.
What determines the invoice is tokens consumed per completed task, and that is a function of how many reasoning tokens the model burns, how many tool calls it makes, how many retries it needs, and how often it goes down a dead end before recovering. A model that costs twice as much per token but finishes in a third of the turns is cheaper.
This was the specific complaint that dogged Fable 5. Business customers and developers criticized the model's burn rate — the volume of tokens it consumed completing tasks — with users exhausting token budgets and running up unexpected bills, according to Fortune's reporting on the launch. Anthropic clearly heard it. Opus 5's entire positioning is a response.
The customer numbers Anthropic published are the most useful evidence here, because they are ratios rather than raw scores:
Harvey reported that Opus 5 maintained comparable performance while generating 26% fewer tokens on average than Opus 4.8 at max reasoning.
A financial modeling team reported 9 percentage points higher accuracy on average across effort levels, with a third fewer turns and tool calls and 60% less time.
A trading firm reported that Opus 5 was the strongest Opus model on their trading benchmark using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8.
Zapier's Wade Foster said the model topped AutomationBench without spending more tokens than prior Claude models, running a full churn-prevention sequence end to end and hitting 100% where previous models did not pass.
A seventh of the reasoning tokens is not an incremental efficiency gain. If that ratio holds outside a single trading benchmark, it changes the arithmetic on workloads that were previously priced out of Opus-tier models entirely.
The caveat: every one of these figures comes from an early-access partner with a commercial relationship to Anthropic and a reason to be quoted favorably. They are useful as directional signals about where the gains cluster — token efficiency, turn count, latency — not as guaranteed outcomes on your workload.
The Effort Ladder and How to Actually Use It
The effort parameter is the operational control that makes the cost story real, and it is widely misunderstood.
Effort governs how much the model thinks, not how much it says. Anthropic's prompting guide is explicit that Opus 5's default user-facing responses run longer than prior Opus models', and lowering effort can reduce thinking volume without reliably shortening the response. If your problem is verbose output, effort is the wrong lever — use explicit length instructions instead.
Practical guidance from the docs and the launch material:
low / medium — routine, well-scoped work. The AutomationBench result suggests low effort on Opus 5 is competitive with other models at their peak, which makes it a genuinely viable production setting rather than a degraded mode.
high — the sensible default for most agentic work.
xhigh / max — capability-sensitive tasks. When running at xhigh or max, set a large max_tokens so the model has room to think and act across subagents and tool calls. 64k is a reasonable starting point.
Two hard behavioral changes to handle before you migrate: thinking is on by default, and passing a disabled-thinking configuration at xhigh or max will fail with a 400. If you are carrying migration code written for Opus 4.6-era manual thinking budgets, it will break.
Verification Behavior Is the Substantive Change
Benchmarks compress. The examples Anthropic published tell you more about what changed than the charts do, and they all point at the same behavior: the model checks its own work before declaring success.
The most striking example involves a Frontier-Bench task where Opus 5 was given a drawing of a machine part and asked to write code rebuilding it as a 3D FreeCAD model — but was intentionally given no way to view the drawing directly. It wrote its own computer vision pipeline to extract the geometry from raw pixels, then reconstructed the part, repeatedly, where no competing model with the same setup succeeded in five attempts.
Two more from the same section: given a real bug in a widely used open-source package manager, Opus 5 found the root cause and fixed an edge case the community's own patch had missed, while a competing model fixed only the surface symptom and reported the bug resolved. And an engineer at a trading firm used it to build a market data feed for a new exchange in one session — with no live feed to validate against, the model built its own test harness to confirm its code parsed the exchange's data correctly.
The common thread is a model that treats "I produced output" and "the output is correct" as different claims. A frontend testing partner described it opening pages in a browser at desktop and phone widths, catching a product hidden below the mobile fold and an off-screen checkout button, and fixing both before handing work back. JetBrains' Denis Shiryaev characterized the difference as judgment — catching logical faults during planning rather than after the fact.
For anyone running long-horizon agents, this is the property that matters most. Autonomous runs fail not because the model cannot solve the task but because it convinces itself it already has. A model that independently constructs verification methods is qualitatively easier to leave unsupervised.
One testimonial is worth noting because it describes something rarer than capability. A partner engineer reported that during a rearchitecting session, Opus 5 pushed back on a proposed design and did not fold when pressed — instead narrowing its objection to a single question and proposing a compromise. Models that cave under user insistence are a well-documented failure mode in agentic settings, and one that is hard to benchmark.
Alignment: The Strongest Claim in the Post
Anthropic states that its automated behavioral audit found Opus 5 to be its most aligned model to date — adhering to Claude's Constitution better than Opus 4.8, Sonnet 5, or Fable 5, exhibiting the lowest rates of deceptive behavior, and being the least susceptible to being tricked into misuse. The reported score is 2.3 on overall misaligned behavior, the lowest of recent models. Anthropic also describes it as its safest model for avoiding reckless actions with hard-to-reverse side effects.
The system card adds that across Usage Policy, user wellbeing, child safety, and bias and integrity evaluations, Opus 5 performs comparably to Opus 4.8, maintaining high harmless-response rates on single-turn harmful requests while keeping among the lowest over-refusal rates on benign requests of any recent model.
That last clause deserves emphasis because over-refusal is the alignment failure that actually costs enterprises money. A model that refuses legitimate security research, medical questions, or financial analysis is not safe — it is unusable for entire departments. Holding refusal rates down while improving harm resistance is the harder engineering problem, and it is the one Anthropic is claiming to have made progress on.
The structural caveat is the same one that applies to every alignment claim in this industry: the audit is largely automated, largely internal, and scored on metrics Anthropic designed. It is more disclosure than most labs provide. It is not independent verification.
The Cyber Split: Finding vs. Exploiting
This is the most technically interesting section of the launch, and the one most likely to be misread.
Anthropic did not train Opus 5 on cyber tasks. As with Opus 4.8, that was deliberate. But the model improved substantially on those tasks anyway as a byproduct of becoming more generally capable, and it now comes close to Mythos 5 at finding cybersecurity vulnerabilities.
The critical distinction: it remains substantially behind Mythos 5 on exploitation — turning vulnerabilities into material cyber threats. Anthropic illustrates this with OSS-Fuzz, an internal evaluation measuring whether models can find and then exploit vulnerabilities without human guidance. Mythos 5 and Opus 5 identify vulnerabilities at similar rates; Opus 5's exploit development score is far behind.
The system card lists a broader battery: ExploitBench, OSS-Fuzz, Firefox 147, plus two newly added benchmarks — CyScenarioBench and ExploitGym — alongside external cyber range testing from the UK AI Security Institute, with the same conclusion: capabilities exceed Opus 4.8, fall short of Mythos 5. The model ships under ASL-3 protections.
That gap between finding and exploiting is doing a lot of load-bearing work in Anthropic's release reasoning. It is a real distinction — vulnerability discovery is closer to code review than to attack — but it is also a gap that general capability gains have been narrowing on their own, without targeted training. Anyone tracking capability thresholds should watch that specific delta across the next two Opus releases rather than the headline scores.
Safeguards That Route Instead of Refuse
The safeguards design is a meaningful departure from how the Fable 5 launch handled the same problem.
On cybersecurity, Opus 5's classifiers are proportionally less restrictive than Fable 5's: they permit finding vulnerabilities in source code, while blocking binary-based vulnerability scanning, penetration testing, and exploit generation. Anthropic expects the classifiers to intervene around 85% less often than on Fable 5.
An 85% reduction in false positives is the single most practically consequential number in the announcement for security teams. Fable 5's conservative tuning was a documented source of friction, and Anthropic acknowledged at that launch that safeguards would sometimes catch harmless requests.
More importantly, the failure mode changed. In Claude.ai, Claude Code, and Claude Cowork, flagged requests fall back to Opus 4.8 by default, and fallbacks can also be enabled on the API. Researchers and enterprises already enrolled in the Cyber Verification Program get immediate access to a version with fewer restrictions.
On biology, the picture inverts. Because Opus 5 carries safeguards similar to Opus 4.8's, it is now Anthropic's most capable generally available model for scientific research — and biology-related requests blocked on Fable 5 now route to Opus 5 rather than Opus 4.8. Anthropic is explicit that the model still shows important limitations on long-running autonomous research tasks, where the most substantial biology-related risks are expected, and Mythos 5 remains stronger for that work.
The design philosophy here is worth naming: degrade rather than deny. A blocked request that silently produces a usable answer from a slightly less capable model is a fundamentally better product experience than a refusal, and it preserves the safety property. This is the architecture the rest of the industry will likely copy.
Two Platform Updates Doing Quiet Heavy Lifting
Alongside the model, Anthropic shipped two beta features that will matter more to agent builders than most of the benchmark chart.
Mid-conversation tool changes. Developers can now change which tools Claude has access to within a conversation without invalidating the prompt cache. If you have built multi-phase agents — a research phase with search tools, a writing phase with file tools, a review phase with test-execution tools — you have been paying a full cache invalidation at every handoff. On a 1M-token context that is a real cost, and it is now gone. Expect this to change how people architect phased agents.
Automatic fallbacks on the API. Requests flagged by safety classifiers on Opus 5 or Fable 5 can now route automatically to another model, so API requests route to the best available model by default rather than being blocked.
The second one is the more significant product decision. It converts safeguards from a hard failure into a graceful degradation, which for anyone running Claude in a production path means the classifier no longer represents an uptime risk. If you have been holding off on Opus-tier models specifically because of unpredictable refusals in automated pipelines, this is the change that addresses it.
If you are running agents in production and want visibility into which model handled which request — including when fallbacks fire — that observability layer is worth building or adopting early. Fallback routing is invisible by default, and a run that silently completed on Opus 4.8 instead of Opus 5 will have different characteristics you will want to see in your traces. Tools like OpenClaw Mission Control exist for exactly this class of problem: watching what your agents actually did rather than what you assumed they did.
Opus 5 or Fable 5: The Actual Decision
Anthropic is not retiring Fable 5, and the distinction it draws is specific.
A company spokesperson told VentureBeat that Fable 5 is intended for the longest, most autonomous jobs, where the model must stay coherent across many connected steps over hours or days with dense source material — and advised customers to test both on a representative workload, one bounded task and one long-horizon job.
That framing is the useful decision rule:
Choose Opus 5 for: bounded tasks with clear success criteria, high-volume production workloads, code review and bug-finding, document and spreadsheet work, cost-sensitive agentic pipelines, computer use, scientific research where Fable's biology classifiers get in the way, and anything where the 30-day retention condition on Fable is a compliance problem.
Choose Fable 5 for: multi-day autonomous runs, work requiring sustained coherence across dense source material, and tasks where a 0.5% capability difference genuinely determines the outcome.
Choose Mythos 5 for: offensive cybersecurity and frontier biology work — if you are one of the small number of organizations with access.
The honest read is that the Opus 5 column covers the overwhelming majority of real production workloads, and the Fable column is narrow and getting narrower. That is presumably intentional. Anthropic's revenue skews heavily toward API and enterprise usage, and a model that captures most of the frontier's value at half the serving cost is better for margins than one that captures all of it at double.
The Competitive Picture
Opus 5 lands into the most crowded frontier the industry has had.
OpenAI's GPT-5.6 Sol launched July 9, replacing GPT-5.5 as ChatGPT's default and taking the top spot on GPQA Diamond, priced at $5 per million input and $30 per million output. Artificial Analysis noted at the time that GPT-5.6 Sol came second to Claude Fable 5 on the Intelligence Index at one third of the cost, and led the Coding Agent Index in OpenAI's Codex harness. Opus 5 now sits above it on the composite index at a lower output price.
Kimi K3 arrived July 16 as the largest open-weight model ever at 2.8T parameters, ranking fourth overall while leading the Frontend Code Arena — which means the open-weight frontier is now roughly four index points behind the closed frontier and undercutting it substantially on price.
Google's position is the wildcard, with Gemini 3.5 Flash and the 3.x Pro line competing primarily on context, speed, and price rather than peak reasoning.
The structural read: general-capability leadership is now measured in single index points and lasts weeks. What does not commoditize as quickly is the operational surface — effort control, cache behavior, fallback routing, subagent coordination, retention terms. Opus 5's launch is disproportionately about that surface, and that is probably the correct strategic bet.
Migration Notes
If you are moving from Opus 4.8 or Opus 4.7, the practical checklist:
Model string — swap to
claude-opus-5. No beta headers.Thinking configuration — thinking is on by default; remove any explicit disable at xhigh or max effort or you will get a 400.
max_tokens — raise it. At xhigh and max, max_tokens is a hard ceiling on thinking plus response text combined. Under-setting it is the most common cause of truncated agentic runs.
Re-tune effort downward, not upward — the consistent signal from partner reports is that Opus 5 at a lower effort setting matches Opus 4.8 at a higher one. Start one rung below where you were and measure.
Response length — expect longer default responses than prior Opus models. Handle this with explicit instructions rather than effort reduction.
Re-baseline your cost model — cost per task, not cost per token. The token-efficiency gains are where the savings live, and they will not show up in a per-token comparison.
Consider enabling automatic fallbacks if you run in a production path where a classifier refusal would break a workflow.
Test subagent workloads specifically — Anthropic highlights improved writer-verifier patterns and fewer cases of agents overwriting each other's work, but also recommends capping delegation on cost-sensitive workloads.
What the Coverage Is Underweighting
Three things the day-one reporting has largely skipped.
The benchmark provenance. Frontier-Bench v0.1 results are from an internal Anthropic run on a specific harness. ARC-AGI 3 at 3× the next-best model is an extraordinary claim published without third-party replication. Both may hold up. Neither should be treated as settled until they do.
The Fable 5 cannibalization. Anthropic released a model that beats its own flagship on the most-cited composite index at half the price, six weeks after that flagship returned from a two-week export-control suspension. Whatever the internal reasoning, the practical effect is that Fable 5's addressable market just narrowed considerably. Customers who onboarded to Fable in June and built around its behavior now have a cheaper option that scores higher on most things they measure.
The cyber trend line. Opus 5 approached Mythos 5 on vulnerability discovery without being trained on cyber tasks at all. That is the finding worth tracking. The exploitation gap is what currently justifies the lighter safeguards, and nothing in the release suggests that gap is structural rather than temporary.
What to Watch Next
Four things over the coming weeks.
Independent benchmark replication. Artificial Analysis has already confirmed the index position at 61. The harder tests are Frontier-Bench and ARC-AGI 3 under third-party harnesses. If the ARC-AGI margin holds outside Anthropic's setup, it is a bigger story than the pricing.
Real-world token burn. Partner reports are unanimous on efficiency gains, but partner reports always are. The signal to watch is developer chatter over the next two weeks about actual bills on actual workloads — that is what exposed Fable 5's burn rate problem, and it will do the same here if there is one.
Whether Fable 5 gets repositioned. Either Anthropic clarifies the long-horizon differentiation with hard evidence, or Fable becomes a legacy tier faster than anyone planned.
The IPO frame. Anthropic filed confidentially in June and has been reported as targeting a listing later this year, against a run-rate revenue figure reported at $47B as of late May. A model that maintains frontier positioning while cutting serving costs is exactly the gross-margin story an S-1 needs. Read the efficiency emphasis in this launch with that in mind — it is genuine engineering, and it is also a narrative being built for a specific audience.
Bottom Line
Claude Opus 5 is not a capability breakthrough. On the composite index it beats Anthropic's own flagship by a single point, which is a tie.
It is a cost breakthrough, and in July 2026 that is the more valuable kind. Frontier-adjacent performance at $5/$25, with a 1M context window, a usable low-effort mode, token consumption reported down sharply against its predecessor, safeguards that degrade instead of refuse, and the lowest misalignment scores Anthropic has published — that combination changes which workloads are economically viable, not just which are technically possible.
The claims deserve independent verification, the customer testimonials are exactly as self-interested as they appear, and the benchmark provenance is thinner than the headlines suggest. But the direction is unambiguous. The frontier is no longer where the competition is happening. Cost per completed task is, and this release is aimed squarely at it.