AIREITER

Claude Sonnet 5 vs Opus 4.8: Which One Should You Use?

Last Updated: 2026-07-01 09:25:02

"Sonnet 5 is close to Opus 4.8, but cheaper" is Anthropic's own pitch. We wanted to know where that stops being true, so we ran four identical tasks through both models via the Claude Code CLI and logged the real cost, duration, and tool-call counts from each API response. The one that mattered most: push Sonnet 5 to xhigh effort to match what Opus 4.8 does by default, and the price gap we measured nearly disappeared — while Opus 4.8 finished in roughly half the time.

Sonnet 5 vs Opus 4.8 at a Glance

Claude Sonnet 5

Claude Opus 4.8

Released

June 30, 2026

May 28, 2026

Context window

1M tokens

1M tokens

Max output

128K tokens

128K tokens

Pricing (in/out per 1M tokens)

$2/$10 intro through Aug 31, 2026, then $3/$15

$5/$25

Fast mode

Not supported

Supported (research preview), up to ~2.5x output speed at premium pricing

Effort levels

low / medium / high / xhigh / max

low / medium / high / xhigh / max

Positioning

Cheaper, agentic-focused Sonnet-tier model

Anthropic's flagship, highest-accuracy option

Context window and max output are identical — not a differentiator here. One thing that is: Sonnet 5 runs on a new tokenizer that Anthropic says produces "approximately 30% more tokens" than Sonnet 4.6 for the same text, so a max_tokens budget or cost estimate carried over from Sonnet 4.6 won't translate directly.

Official Benchmark Comparison

Per Anthropic's own disclosures, compiled alongside third-party analysis by llm-stats.com, the pattern is consistent: Sonnet 5 wins or ties a couple of benchmarks, and Opus 4.8 leads on most others, usually by single digits.

Benchmark

Sonnet 5

Opus 4.8

Gap

Terminal-Bench 2.1

80.4%

74.6%

Sonnet 5 +5.8

Humanity's Last Exam (with tools)

57.4%

57.9%

Near tie

Humanity's Last Exam (no tools)

43.2%

49.8%

Opus +6.6

SWE-bench Verified

85.2%

88.6%

Opus +3.4

SWE-bench Pro

63.2%

69.2%

Opus +6.0

Toolathlon

54.3%

59.9%

Opus +5.6

OSWorld-Verified (computer use)

81.2%

83.4%

Opus +2.2

CursorBench

61.2%

63.8%

Opus +2.6

USAMO 2026 problems

79.5%

96.7%

Opus +17.2

Two things stand out. Opus 4.8's biggest lead is on hard math (USAMO), not coding — the coding gaps (SWE-bench, Terminal-Bench) are all single digits, and Sonnet 5 actually wins one of them outright. On tool-augmented general reasoning (HLE with tools), the two are close enough to call it a tie.

There's also a safety number worth flagging even though it's not a capability benchmark: per the same disclosures, in browser-use scenarios without additional safeguards, Sonnet 5's measured prompt-injection attack success rate was 0.93%, against 31.5% for Opus 4.8. Anthropic hasn't published a detailed explanation for that specific gap. If you're building anything agentic that browses the open web unsupervised, it's worth testing your own safeguards rather than assuming the flagship model is automatically the safer default.

We Ran Four Head-to-Head Tests Ourselves

Most of what's out there right now either compiles the official benchmark table above or shares subjective, single-model impressions. We didn't find a controlled, same-prompt cost-and-latency comparison, so we ran our own.

Methodology: four tasks, run on 2026-07-01, same prompt sent to claude-sonnet-5 and claude-opus-4-8 each time via the Claude Code CLI (claude -p --model <id> --effort <level> --output-format json). Cost, duration, and turn counts are read directly from each run's API response, not estimated from token counts. Each model/effort configuration was run once per task — real numbers from a single run, not a statistically averaged sample. Treat the size of each gap as directional rather than exact; the direction was consistent across every test we ran.

Test 1: The Cost-Reversal Check

The question we most wanted to answer: does Sonnet 5 stay cheaper once you push its effort level up to compensate for a harder task? We ran the same coding task as Test 2 (below) at Sonnet 5 xhigh against Opus 4.8 medium.

Setup

Cost

Duration

Turns

Result

Sonnet 5, effort xhigh

$0.390

40.5s

7

8/8 tests pass

Opus 4.8, effort medium

$0.401

22.9s

4

7/7 tests pass

A 3% cost difference — essentially a wash — and Opus 4.8 on its lower effort setting finished in 57% of the time using almost half the turns. Both produced correct, working code. This lines up with what people are already flagging on Hacker News: one commenter there estimated Opus 4.8 costing roughly $0.45 on medium reasoning against Sonnet 5 at roughly $0.52 on xhigh/max for comparable work.

Terminal output showing the actual `claude -p` calls for the Sonnet 5 xhigh vs Opus 4.8 medium test, with the real duration_ms, num_turns, and total_cost_usd fields from each API response

Test 2: Coding Task at Matched Effort

Same prompt, both models at high effort: write an efficient longest-palindromic-substring function, generate test cases covering the edge cases, run them, fix anything that fails.

Setup

Cost

Duration

Turns

Result

Sonnet 5 (high)

$0.378

37.4s

7

8/8 pass

Opus 4.8 (high)

$0.439

28.9s

5

8/8 pass

Both landed on the same expand-around-center approach and passed every test on the first run. Opus 4.8 got there faster and in fewer turns, for about 16% more money. When quality is identical, this one goes to Opus on speed alone.

Test 3: Writing / Knowledge Work

A business-judgment prompt with no single correct answer: advise a 12-person SaaS company on whether to spend 3-4 weeks migrating to a second AWS region for disaster recovery, six weeks before their Series A closes.

Setup

Cost

Duration

Turns

Sonnet 5 (high)

$0.072

14.7s

1

Opus 4.8 (high)

$0.084

22.7s

1

Both gave essentially the same recommendation — skip the full migration, ship a lightweight backup/runbook instead — with comparable reasoning quality. Sonnet 5 got there in about two-thirds the time and cost 14% less. This is the one test where "Sonnet 5 holds up fine on knowledge work" came through cleanly.

Test 4: Agentic Lookup — Does Sonnet 5 Really "Overthink"?

Some Reddit threads describe Sonnet 5 as more prone to overthinking simple requests than Opus 4.8. We tested a genuinely simple one: sum the byte size of every .py file in a directory tree and report which subdirectory has the most of them.

Setup

Cost

Duration

Turns

Sonnet 5 (high)

$0.119

18.5s

3

Opus 4.8 (high)

$0.120

12.1s

2

Both landed on the identical correct answer. Cost was a rounding error apart, but Sonnet 5 took one extra tool-call turn and about 50% longer. That's a real, if modest, data point for the "overthinking" complaint — and it lines up with Test 1: Sonnet 5 tends to take more steps to arrive at the same place.

So Which Should You Actually Use?

  • Pick Opus 4.8 for coding tasks where speed matters, agentic workflows with many tool calls, or anything where you'd otherwise be tempted to push Sonnet 5's effort past high to trust the output — that's exactly the situation where Sonnet 5's cost advantage disappears.

  • Pick Sonnet 5 for high-volume, budget-sensitive work and general knowledge/writing tasks, but keep it at medium or high effort. Don't reflexively bump it to xhigh "just to be safe" — that's where the pricing story falls apart.

  • Either works for simple lookups and single-turn requests; the practical difference we measured was one extra turn and a few seconds, not a wrong answer.

FAQ

Is Claude Sonnet 5 actually cheaper than Opus 4.8?

At matched effort levels, yes — noticeably. But push Sonnet 5 to xhigh to compensate for a harder task and the gap can shrink to a few percent, while Opus 4.8 on a lower effort setting finishes faster. Check what effort level you're actually running before assuming you're saving money. For the full price list across Haiku, Sonnet, and Opus, see our Claude API pricing guide.

What about Claude Sonnet 5 vs Sonnet 4.6?

A generational upgrade, not a model-tier choice — and cost doesn't move the same direction there either. Sonnet 5 beats Sonnet 4.6 by double digits on Terminal-Bench, but in our own head-to-head cost tests, Sonnet 5 came out more expensive than Sonnet 4.6 at every effort level we tried, mainly due to the new tokenizer. Full tests and numbers in Sonnet 5 vs Sonnet 4.6: Is It Actually Cheaper?

Is Claude Sonnet 5 vs Opus 4.6 a fair comparison?

Not really — Opus 4.6 is a generation behind Opus 4.8. If you're deciding what to use today, compare against Opus 4.8, not its predecessor.

Does the context window differ between the two?

No — both offer a 1M-token context window and 128K max output.

Why is Opus 4.8's prompt-injection rate so much higher in browser use?

Per Anthropic's disclosures, 31.5% without additional safeguards against Sonnet 5's 0.93%, specifically in unsupervised browser-use scenarios. Anthropic hasn't published a detailed explanation. Treat it as a reason to test your specific safeguards rather than assuming the flagship model is automatically the safer default.

Appendix: Exact Prompts and Raw Output

For anyone who wants to reproduce this, here are the exact prompts we sent for each test (Test 1 and Test 2 used the same coding prompt, just at different effort levels).

Tests 1 & 2 — coding prompt:

Write a Python function `longest_palindromic_substring(s: str) -> str` that returns
the longest palindromic substring of s, using an approach more efficient than
brute-force O(n^3) (e.g. expand-around-center or Manacher's algorithm). Save it
to solution.py.

Then write test_solution.py with at least 6 test cases covering: empty string,
single character, all-same characters, no palindrome longer than 1, an
even-length palindrome, and an odd-length palindrome.

Run the tests with pytest and make sure they all pass. If any fail, fix the
code and re-run until all pass. Report the final pytest output.

Test 3 — writing/knowledge-work prompt:

You're advising a small SaaS company (12 employees, $80k MRR, currently on a
single AWS region) on whether to expand to a second AWS region for disaster
recovery before their Series A fundraising, which closes in 6 weeks.
Engineering estimates the migration takes 3-4 weeks and would consume most of
the team's capacity during that window, delaying two customer-requested
features.

Write a ~350-word executive recommendation: should they do it now, delay it,
or find a middle path? Justify with tradeoffs. No preamble, just the
recommendation.

Test 4 — agentic lookup prompt:

In the current directory tree, find all files with a .py extension, sum their
total size in bytes, and tell me which single subdirectory (immediate child of
the working directory) contains the most .py files by count. Just the two
numbers/answer, no need to write any new files.

Here's the actual API response for the Sonnet 5 (xhigh) run from Test 1, trimmed to the fields relevant to cost and timing:

{
  "duration_ms": 40474,
  "num_turns": 7,
  "stop_reason": "end_turn",
  "total_cost_usd": 0.38962215,
  "usage": {
    "input_tokens": 4303,
    "cache_creation_input_tokens": 65277,
    "cache_read_input_tokens": 310598,
    "output_tokens": 2583
  }
}

That's the unedited total_cost_usd, duration_ms, and token counts the Claude Code CLI returned for that run — the numbers in the Test 1 table above come directly from fields like these, one JSON response per run.