"Sonnet 5 is close to Opus 4.8, but cheaper" is Anthropic's own pitch. We wanted to know where that stops being true, so we ran four identical tasks through both models via the Claude Code CLI and logged the real cost, duration, and tool-call counts from each API response. The one that mattered most: push Sonnet 5 to xhigh effort to match what Opus 4.8 does by default, and the price gap we measured nearly disappeared — while Opus 4.8 finished in roughly half the time.
Sonnet 5 vs Opus 4.8 at a Glance
Claude Sonnet 5 | Claude Opus 4.8 | |
|---|---|---|
Released | June 30, 2026 | |
Context window | 1M tokens | 1M tokens |
Max output | 128K tokens | 128K tokens |
Pricing (in/out per 1M tokens) | $5/$25 | |
Fast mode | Not supported | Supported (research preview), up to ~2.5x output speed at premium pricing |
Effort levels | low / medium / high / xhigh / max | low / medium / high / xhigh / max |
Positioning | Cheaper, agentic-focused Sonnet-tier model | Anthropic's flagship, highest-accuracy option |
Context window and max output are identical — not a differentiator here. One thing that is: Sonnet 5 runs on a new tokenizer that Anthropic says produces "approximately 30% more tokens" than Sonnet 4.6 for the same text, so a max_tokens budget or cost estimate carried over from Sonnet 4.6 won't translate directly.
Official Benchmark Comparison
Per Anthropic's own disclosures, compiled alongside third-party analysis by llm-stats.com, the pattern is consistent: Sonnet 5 wins or ties a couple of benchmarks, and Opus 4.8 leads on most others, usually by single digits.
Benchmark | Sonnet 5 | Opus 4.8 | Gap |
|---|---|---|---|
Terminal-Bench 2.1 | 80.4% | 74.6% | Sonnet 5 +5.8 |
Humanity's Last Exam (with tools) | 57.4% | 57.9% | Near tie |
Humanity's Last Exam (no tools) | 43.2% | 49.8% | Opus +6.6 |
SWE-bench Verified | 85.2% | 88.6% | Opus +3.4 |
SWE-bench Pro | 63.2% | 69.2% | Opus +6.0 |
Toolathlon | 54.3% | 59.9% | Opus +5.6 |
OSWorld-Verified (computer use) | 81.2% | 83.4% | Opus +2.2 |
CursorBench | 61.2% | 63.8% | Opus +2.6 |
USAMO 2026 problems | 79.5% | 96.7% | Opus +17.2 |
Two things stand out. Opus 4.8's biggest lead is on hard math (USAMO), not coding — the coding gaps (SWE-bench, Terminal-Bench) are all single digits, and Sonnet 5 actually wins one of them outright. On tool-augmented general reasoning (HLE with tools), the two are close enough to call it a tie.
There's also a safety number worth flagging even though it's not a capability benchmark: per the same disclosures, in browser-use scenarios without additional safeguards, Sonnet 5's measured prompt-injection attack success rate was 0.93%, against 31.5% for Opus 4.8. Anthropic hasn't published a detailed explanation for that specific gap. If you're building anything agentic that browses the open web unsupervised, it's worth testing your own safeguards rather than assuming the flagship model is automatically the safer default.
We Ran Four Head-to-Head Tests Ourselves
Most of what's out there right now either compiles the official benchmark table above or shares subjective, single-model impressions. We didn't find a controlled, same-prompt cost-and-latency comparison, so we ran our own.
Methodology: four tasks, run on 2026-07-01, same prompt sent to claude-sonnet-5 and claude-opus-4-8 each time via the Claude Code CLI (claude -p --model <id> --effort <level> --output-format json). Cost, duration, and turn counts are read directly from each run's API response, not estimated from token counts. Each model/effort configuration was run once per task — real numbers from a single run, not a statistically averaged sample. Treat the size of each gap as directional rather than exact; the direction was consistent across every test we ran.
Test 1: The Cost-Reversal Check
The question we most wanted to answer: does Sonnet 5 stay cheaper once you push its effort level up to compensate for a harder task? We ran the same coding task as Test 2 (below) at Sonnet 5 xhigh against Opus 4.8 medium.
Setup | Cost | Duration | Turns | Result |
|---|---|---|---|---|
Sonnet 5, effort | $0.390 | 40.5s | 7 | 8/8 tests pass |
Opus 4.8, effort | $0.401 | 22.9s | 4 | 7/7 tests pass |
A 3% cost difference — essentially a wash — and Opus 4.8 on its lower effort setting finished in 57% of the time using almost half the turns. Both produced correct, working code. This lines up with what people are already flagging on Hacker News: one commenter there estimated Opus 4.8 costing roughly $0.45 on medium reasoning against Sonnet 5 at roughly $0.52 on xhigh/max for comparable work.

Test 2: Coding Task at Matched Effort
Same prompt, both models at high effort: write an efficient longest-palindromic-substring function, generate test cases covering the edge cases, run them, fix anything that fails.
Setup | Cost | Duration | Turns | Result |
|---|---|---|---|---|
Sonnet 5 (high) | $0.378 | 37.4s | 7 | 8/8 pass |
Opus 4.8 (high) | $0.439 | 28.9s | 5 | 8/8 pass |
Both landed on the same expand-around-center approach and passed every test on the first run. Opus 4.8 got there faster and in fewer turns, for about 16% more money. When quality is identical, this one goes to Opus on speed alone.
Test 3: Writing / Knowledge Work
A business-judgment prompt with no single correct answer: advise a 12-person SaaS company on whether to spend 3-4 weeks migrating to a second AWS region for disaster recovery, six weeks before their Series A closes.
Setup | Cost | Duration | Turns |
|---|---|---|---|
Sonnet 5 (high) | $0.072 | 14.7s | 1 |
Opus 4.8 (high) | $0.084 | 22.7s | 1 |
Both gave essentially the same recommendation — skip the full migration, ship a lightweight backup/runbook instead — with comparable reasoning quality. Sonnet 5 got there in about two-thirds the time and cost 14% less. This is the one test where "Sonnet 5 holds up fine on knowledge work" came through cleanly.
Test 4: Agentic Lookup — Does Sonnet 5 Really "Overthink"?
Some Reddit threads describe Sonnet 5 as more prone to overthinking simple requests than Opus 4.8. We tested a genuinely simple one: sum the byte size of every .py file in a directory tree and report which subdirectory has the most of them.
Setup | Cost | Duration | Turns |
|---|---|---|---|
Sonnet 5 (high) | $0.119 | 18.5s | 3 |
Opus 4.8 (high) | $0.120 | 12.1s | 2 |
Both landed on the identical correct answer. Cost was a rounding error apart, but Sonnet 5 took one extra tool-call turn and about 50% longer. That's a real, if modest, data point for the "overthinking" complaint — and it lines up with Test 1: Sonnet 5 tends to take more steps to arrive at the same place.
So Which Should You Actually Use?
Pick Opus 4.8 for coding tasks where speed matters, agentic workflows with many tool calls, or anything where you'd otherwise be tempted to push Sonnet 5's effort past
highto trust the output — that's exactly the situation where Sonnet 5's cost advantage disappears.Pick Sonnet 5 for high-volume, budget-sensitive work and general knowledge/writing tasks, but keep it at
mediumorhigheffort. Don't reflexively bump it toxhigh"just to be safe" — that's where the pricing story falls apart.Either works for simple lookups and single-turn requests; the practical difference we measured was one extra turn and a few seconds, not a wrong answer.
FAQ
Is Claude Sonnet 5 actually cheaper than Opus 4.8?
At matched effort levels, yes — noticeably. But push Sonnet 5 to xhigh to compensate for a harder task and the gap can shrink to a few percent, while Opus 4.8 on a lower effort setting finishes faster. Check what effort level you're actually running before assuming you're saving money. For the full price list across Haiku, Sonnet, and Opus, see our Claude API pricing guide.
What about Claude Sonnet 5 vs Sonnet 4.6?
A generational upgrade, not a model-tier choice — and cost doesn't move the same direction there either. Sonnet 5 beats Sonnet 4.6 by double digits on Terminal-Bench, but in our own head-to-head cost tests, Sonnet 5 came out more expensive than Sonnet 4.6 at every effort level we tried, mainly due to the new tokenizer. Full tests and numbers in Sonnet 5 vs Sonnet 4.6: Is It Actually Cheaper?
Is Claude Sonnet 5 vs Opus 4.6 a fair comparison?
Not really — Opus 4.6 is a generation behind Opus 4.8. If you're deciding what to use today, compare against Opus 4.8, not its predecessor.
Does the context window differ between the two?
No — both offer a 1M-token context window and 128K max output.
Why is Opus 4.8's prompt-injection rate so much higher in browser use?
Per Anthropic's disclosures, 31.5% without additional safeguards against Sonnet 5's 0.93%, specifically in unsupervised browser-use scenarios. Anthropic hasn't published a detailed explanation. Treat it as a reason to test your specific safeguards rather than assuming the flagship model is automatically the safer default.
Appendix: Exact Prompts and Raw Output
For anyone who wants to reproduce this, here are the exact prompts we sent for each test (Test 1 and Test 2 used the same coding prompt, just at different effort levels).
Tests 1 & 2 — coding prompt:
Write a Python function `longest_palindromic_substring(s: str) -> str` that returns
the longest palindromic substring of s, using an approach more efficient than
brute-force O(n^3) (e.g. expand-around-center or Manacher's algorithm). Save it
to solution.py.
Then write test_solution.py with at least 6 test cases covering: empty string,
single character, all-same characters, no palindrome longer than 1, an
even-length palindrome, and an odd-length palindrome.
Run the tests with pytest and make sure they all pass. If any fail, fix the
code and re-run until all pass. Report the final pytest output.
Test 3 — writing/knowledge-work prompt:
You're advising a small SaaS company (12 employees, $80k MRR, currently on a
single AWS region) on whether to expand to a second AWS region for disaster
recovery before their Series A fundraising, which closes in 6 weeks.
Engineering estimates the migration takes 3-4 weeks and would consume most of
the team's capacity during that window, delaying two customer-requested
features.
Write a ~350-word executive recommendation: should they do it now, delay it,
or find a middle path? Justify with tradeoffs. No preamble, just the
recommendation.
Test 4 — agentic lookup prompt:
In the current directory tree, find all files with a .py extension, sum their
total size in bytes, and tell me which single subdirectory (immediate child of
the working directory) contains the most .py files by count. Just the two
numbers/answer, no need to write any new files.
Here's the actual API response for the Sonnet 5 (xhigh) run from Test 1, trimmed to the fields relevant to cost and timing:
{
"duration_ms": 40474,
"num_turns": 7,
"stop_reason": "end_turn",
"total_cost_usd": 0.38962215,
"usage": {
"input_tokens": 4303,
"cache_creation_input_tokens": 65277,
"cache_read_input_tokens": 310598,
"output_tokens": 2583
}
}
That's the unedited total_cost_usd, duration_ms, and token counts the Claude Code CLI returned for that run — the numbers in the Test 1 table above come directly from fields like these, one JSON response per run.