Claude Opus 5 vs Sonnet 5: Who Wins at the Same Budget

Last Updated: 2026-07-25 05:04:31

Sonnet 5 shipped on June 30, 2026, and r/ClaudeAI immediately produced a complaint its price sheet does not predict: "Sonnet 5 is worse than Opus at the same price at high and xhigh?"

One poster reduced the mechanism to a sentence. Sonnet is "actually cheaper per token, but it takes way more steps and tokens to finish the tasks, so that's what makes it lose its price advantage."

They were right, with one exception. Claude Opus 5 vs Sonnet 5 at a matched budget favours Opus nearly everywhere: above about $2.30 per task, Opus one effort tier down beats Sonnet one tier up on both cost and quality. Below that line, and on short single-shot work, Sonnet 5 is cheaper by 3.4x.

What each model wins

Claude Opus 5Claude Sonnet 5
Model IDclaude-opus-5claude-sonnet-5
ShippedJuly 24, 2026June 30, 2026
Input / output per MTok$5 / $25$2 / $10 to Aug 31, 2026, then $3 / $15
Batch input / output$2.50 / $12.50$1 / $5 to Aug 31, 2026
Effort tierslow, medium, high, xhigh, max (default high)the same five tiers (default high)
Extended thinkingon by defaultadaptive
Context / max output1M / 128k tokens1M / 128k tokens
Knowledge cutoffMay 2026January 2026
Per-token discountnone60% off Opus today, 40% off from Sept 1
Strongest fitlong-horizon agentic runs, hard reasoning, work that is expensive to reviewsingle-shot extraction, high concurrency, latency-sensitive UX
Our four-task smoke test4/4 passed, 1,182 output tokens, 2.955¢4/4 passed, 858 output tokens, 0.858¢

No public long-horizon benchmark has run Opus 5 yet, so Opus 4.8 holds the Opus column in the per-task tables below. Treat those rows as a conservative floor: the two models are priced identically, and Anthropic puts Opus 5 ahead of 4.8 on agentic work.

A budget-matched Claude Opus 5 vs Sonnet 5 comparison works at all because both models take the same effort parameter across the same five tiers with no beta header, so capability is bought by moving the dial rather than by switching models. Prices are from the Claude Platform pricing page, verified July 25, 2026.

Anthropic's Claude Platform Docs page for the effort parameter, listing Claude Fable 5, Claude Mythos 5, Claude Opus 5, Claude Opus 4.8, Claude Sonnet 5 and others as supported models, and noting that Claude uses high effort by default

Tier by tier: where the two cost curves cross

The per-tier numbers come from DeepSWE v1.1, Datacurve's long-horizon agentic leaderboard:

  • 113 tasks written from scratch, not lifted from merged PRs, which is the basis for its contamination-free claim
  • 91 repositories, 5 languages
  • One shared harness for every model, four runs per task
  • Last updated July 21, three days before Opus 5 shipped

So Opus 4.8 sits in the Opus column, and the one open variable is whether Opus 5 spends more tokens per task than 4.8 did.

EffortSonnet 5 $/taskOpus 4.8 $/taskSonnet ÷ OpusSonnet pass@1Opus pass@1
low$2.19$2.290.95x30.5%40.8%
medium$4.08$3.441.18x39.8%48.7%
high$7.43$4.281.73x48.2%51.8%
xhigh$11.89$8.011.49x49.7%54.4%
max$26.40$13.222.00x53.8%59.0%

Sonnet 5 comes out cheaper at exactly one tier, low, and that is the tier where it gives up 10.3 points of pass rate.

Why the cheaper model bills more

Step count, not verbosity.

EffortSonnet 5 median stepsOpus 4.8 median stepsSonnet input per taskOpus input per task
low70474.5M2.4M
max26011672.4M17.0M

A 60% per-token discount cannot survive reading four times as much text.

Why the dollar column is trustworthy

One caveat first: DeepSWE does not publish its price basis.

A separate 24-task run from July 13 covered both models at all five tiers inside Claude Code. Scale DeepSWE's Sonnet column by 0.67, introductory $2/$10 against standard $3/$15, and the two sets agree within 0.03 at every tier.

Two independent task sets, one crossover, in the same place between medium and high. The agreement is also good evidence DeepSWE billed Sonnet at standard rates.

Its author builds and sells Stet, the harness linked from that post, so treat it as a vendor's own test.

Same budget, different tiers: the comparison that decides it

Matched-tier tables answer the wrong question, because nobody buys an effort tier. You buy an outcome at a price, and pricing Claude Opus 5 vs Sonnet 5 that way reorders the comparison.

Below, for each Sonnet 5 configuration, the cheapest Opus 4.8 tier that matches or beats its pass rate. The intro column is the standard-rate figure times 0.67.

Sonnet 5 configSonnet pass@1Sonnet $, standardSonnet $, intro (est.)Cheapest Opus 4.8 tier that matches itOpus $Opus pass@1Opus cost vs intro SonnetQuality delta
low30.5%$2.19$1.47low$2.2940.8%1.56x+10.3pp
medium39.8%$4.08$2.73low$2.2940.8%0.84x+1.0pp
high48.2%$7.43$4.98medium$3.4448.7%0.69x+0.4pp
xhigh49.7%$11.89$7.97high$4.2851.8%0.54x+2.1pp
max53.8%$26.40$17.69xhigh$8.0154.4%0.45x+0.5pp

Read down the last two columns. Dropping Opus one tier below the Sonnet tier you were considering buys the same pass rate or better for 16% to 55% less, even against promotional pricing. Only the floor row inverts: at low effort promotional Sonnet 5 is 36% cheaper, and hands back 10.3 points of pass rate.

The reading that goes the other way

A commenter in that first thread took the official chart at medium effort: Opus near $7 against Sonnet's $4.50 for success rates of 68% and 62%, "a 55% increase in price for less than a 10% boost in success rate on this test."

That trade is reasonable where a 6-point miss costs nothing, and wrong where a bad diff reaches production.

Both readings sit on someone else's task distribution. Twenty to fifty tasks of your own, scored on cost per successful completion rather than cost per token, beat any table on this page.

What Opus 5 changes at the same $5 and $25

Opus 5's launch pricing matches Opus 4.8 line for line, so the open question is the third axis: tokens per task. The evidence points both ways.

Points to Opus 5 costing less per taskPoints to it costing more
Anthropic reports Opus 5 topping Frontier-Bench v0.1 and "more than doubles Opus 4.8's performance at a lower cost per task" (vendor-reported, its own runs, no independent reproduction)Extended thinking is on by default, where Opus 4.8 had it off
An unnamed legal partner saw roughly 26% fewer tokens at max reasoning (relayed by Anthropic, no methodology given)Our four tasks: Opus 4.8 spent 380 output tokens, Opus 5 spent 1,182

Either direction is live, so re-measure the cost columns on your own repository before committing a budget.

One trap for anyone cutting Opus 5 cost the obvious way: thinking: {"type": "disabled"} is accepted only at effort high and below and returns a 400 at xhigh or max, one of several breaking changes between Opus 4.8 and Opus 5 worth reading before a migration.

Where Sonnet 5 wins: short, single-shot work

We ran both models through four small tasks on July 24, 2026, each graded by script rather than opinion: a two-defect bug fix in kth_smallest, a strict JSON emission with fixed key order and a self-consistent checksum, a Frobenius-number question answering 65, and a refactor adding validation while keeping the signature and carrying zero comments.

Both passed all four. The difference showed up on the meter.

Claude Sonnet 5Claude Opus 5
Checks passed4/44/4
Output tokens, four tasks8581,182
Total latency13.7s24.6s
Output cost at intro / standard rates0.858¢ / 1.287¢2.955¢
Horizontal bar chart of output-token cost in US cents for the same four tasks: Sonnet 5 at introductory $10 per MTok costs 0.858 cents, Sonnet 5 at standard $15 per MTok costs 1.287 cents, and Opus 5 at $25 per MTok costs 2.955 cents

Identical verified outputs, 3.4x the cost on Opus 5 at today's Sonnet pricing, or 2.3x once the promotion ends, in a little over half the wall-clock time. Repeating the two coding tasks through Claude Code's first-party channel pointed the same way: 240 output tokens for Sonnet 5 against 465 for Opus 5.

Three limits matter more than that number:

  • One run per model per task. A smoke test, not a benchmark.
  • Output-token costs only, at official output rates, because the gateway injected a variable system prompt that left input-token counts unusable for billing math.
  • Four short prompts, which is the one regime the community argument excludes. Nothing here runs 70 steps, so nothing here reproduces the step blowup.

Inside that boundary Sonnet 5's discount survives intact. One developer described exactly that shape of work: document-heavy agents doing "single-shot extractions with minimal reasoning needed", where "Opus would be overkill."

Two dates attached to every number here

June 30, 2026: the official cost chart was revised

Anthropic's Claude Sonnet 5 announcement carries a changelog note. The original cost-performance chart used "a simpler methodology that did not reflect the standard methodology we use for agentic search evaluations," which had "the result of underestimating Sonnet 5's performance on the evaluation."

The chart and surrounding text were revised to match the system card's BrowseComp methodology, "which used a 10M token budget with compaction and programmatic tool calling."

So if you are quoting that chart, check which revision you have and read the caption: it prices Sonnet 5 at standard $3/$15, notes introductory pricing puts real-world cost below the chart, and prices Opus 4.8 at $5/$25.

September 1, 2026: the promotional pricing ends

Sonnet 5's introductory rate expires, every published line rises 50%, and Opus 5 stays at $5/$25. That leaves 37 days of runway from publication, after which any spend model built on today's rates is wrong.

One more date gets cited in these arguments without applying here. The Claude 4.7 tokenizer produces roughly 30% more tokens for the same text than Sonnet 4.6's did, so pricing Sonnet 5 against Sonnet 4.6 per token understates the real increase. Opus 5 and Sonnet 5 share that tokenizer, so the tax cancels between them.

Which to pick, by budget

Claude Opus 5 vs Sonnet 5 resolves into three routing rules rather than one winner. The thresholds are DeepSWE's mean cost per task with Opus 4.8 in the Opus seat, so they are opening brackets for repository work, not constants for research agents, support pipelines, or document workflows.

  • Below roughly $2.30 per task, take Sonnet 5 at low effort. The one configuration where its price advantage holds up on agentic work, at the cost of a 30.5% pass rate. Single-shot extraction, classification, anything with a bounded output: stay here and pocket the 3.4x we measured. One developer called Sonnet 5 "close enough to Opus for 90% of the work, but instant (comparatively)", which matters when a human is waiting.
  • From about $2.30 to $8 per task, take Opus one tier down instead of Sonnet one tier up. Opus medium as a working default, Opus high when a mistake is expensive to catch in review. On Opus 5 the price per token is identical but the tokens per task are unmeasured, so treat the bracket as where to start testing.
  • Above $8, stop climbing the Sonnet dial. Sonnet 5 at max cost $26.40 per task for 53.8% on DeepSWE; Opus 4.8 at xhigh cost $8.01 for 54.4%; Fable 5 at low effort reached 59.6% for $3.76. AIReiter lists Opus 4.8 at $1.56/$7.76 per MTok, roughly 69% under list, though its Claude lineup does not yet carry Opus 5.

FAQ

Is Claude Opus 5 more expensive than Sonnet 5?

Per token, yes: $5/$25 against Sonnet 5's promotional $2/$10, so 2.5x the price until August 31, 2026 and 1.67x after. Per finished task the ordering can invert, and medians confirm it rather than a few runaway runs. On DeepSWE the median cost at max effort was $23.28 for Sonnet 5 against $12.35 for Opus 4.8.

How much does Claude Sonnet 5 cost?

$2 per million input tokens and $10 per million output through August 31, 2026, then $3 and $15 from September 1. That rise is 50% across every published line: 5-minute cache writes go $2.50 to $3.75, 1-hour cache writes $4 to $6, cache hits $0.20 to $0.30. Batch pricing is $1/$5 during the promotion.

Why does a cheaper model end up costing more?

Because the bill scales with steps, not with tasks. Across the effort dial Sonnet 5's median step count climbs from 70 to 260 against Opus 4.8's 47 to 116, and its input reading reaches 72.4M tokens per task against 17.0M at max. Four times the reading swamps a 60% discount on each token of it.

Is Opus 5 much better than Sonnet 5?

On single attempts, yes: Opus 4.8 led Sonnet 5 by 3.5 to 10.3 points of pass@1 on DeepSWE, widest at the low end.

Give both four attempts and the gap nearly closes: Sonnet 5 at high effort reached 79.6% pass@4 against Opus 4.8's 77.9%, and at max the two sat at 78.8% and 79.3%. It can get there; it needs more tries, and you pay for each one.

Does Sonnet 5's September 1 price rise change the answer?

It moves the crossover down rather than up. Sonnet 5's per-token advantage narrows from 60% to 40% while Opus 5 holds at $5/$25, so every pairing in the budget-matched table gets worse for Sonnet. The one row that favoured Sonnet, its low tier at 36% below Opus low on introductory pricing, shrinks to roughly 4% below at standard rates.

Related reading