AIREITER

Claude Sonnet 5 vs Sonnet 4.6: Is It Actually Cheaper?

Last Updated: 2026-07-01 09:18:28

Sonnet 5's headline price is lower than Sonnet 4.6's ($2/$10 introductory through Aug 31, 2026, vs. Sonnet 4.6's $3/$15), and Anthropic's own launch chart calls it a "strict improvement" over Sonnet 4.6. We ran both models on the same task at several effort levels to see what that means in practice. Result: on our test, Sonnet 5 cost more than Sonnet 4.6 at every effort level we tried — not less. Here's what we found and why.

One task, single runs. Everything below comes from one coding task, one run per model/effort configuration. Treat the specific numbers as directional, not a statistically averaged benchmark — the exact prompt and a sample of the raw API output are in the appendix if you want to reproduce it on your own workload.

Is Sonnet 5 Actually Cheaper Than Sonnet 4.6? We Tested It.

We gave both models the same task — build a token-bucket rate limiter in Python with tests, run the tests, fix anything that fails — via the Claude Code CLI (claude -p --model <id> --effort <level> --output-format json), so cost and duration come straight from the API response. Run on 2026-07-01.

Setup

Cost

Duration

Turns

Result

Sonnet 5, effort medium

$0.344

31.6s

5

6/6 tests pass

Sonnet 4.6, effort low

$0.261

36.0s

6

7/7 tests pass

Sonnet 4.6, effort medium

$0.253

35.3s

5

7/7 tests pass

Sonnet 5, effort high

$0.349

36.4s

5

8/8 tests pass

Sonnet 5 was faster at medium effort, but cost more than Sonnet 4.6 at both of the Sonnet 4.6 configurations we tested — including Sonnet 4.6 running at a lower effort setting. This is the opposite of some anecdotal community reports claiming Sonnet 5 gets cheaper than Sonnet 4.6 once you compare matched effort tiers.

The likely cause: Sonnet 5 runs on a new tokenizer. Anthropic's own documentation states it produces "approximately 30% more tokens" than Sonnet 4.6 for the same text, and independent testing by Simon Willison found English content specifically running up to ~1.4x more tokens. Sonnet 5's lower sticker price didn't fully offset that on our test task, and after August 31, 2026 — when Sonnet 5's introductory pricing expires and both models sit at the identical $3/$15 sticker rate — the tokenizer difference alone suggests Sonnet 5 is likely to cost more for equivalent English-language work, not the same, absent any other pricing changes. (Full breakdown by language in our Sonnet 5 pricing guide.)

What Else Changed, Beyond Cost

On official benchmarks, Sonnet 5 posts real gains: 80.4% vs. 67.0% on Terminal-Bench 2.1, 63.2% vs. 58.1% on SWE-bench Pro, both per Anthropic's Claude Sonnet 5 system card (see our full benchmark table including Opus 4.8). Anthropic's own launch announcement chart describes Sonnet 5 as a "strict improvement" over Sonnet 4.6.

Looking at our own four test outputs, one behavioral difference stood out that the benchmark table doesn't capture: Sonnet 5 added things we didn't ask for. At both effort levels, it used a monkeypatched fake clock in the tests to avoid real time.sleep() calls, and at high effort it added thread-safety to the rate limiter without being asked. Sonnet 4.6, at both effort levels, stayed closer to the literal task description — using real time.sleep() in tests, and at medium effort it added a summary table for readability but nothing to the actual implementation beyond what was requested.

That lines up with real user feedback posted within the first day of Sonnet 5's launch: one r/claude thread described it as "an objective upgrade from Sonnet 4.6 in terms of reasoning... but it also feels WAY more on edge than 4.6." Our test is a sample of one task, but "adds unrequested scope more readily" is a concrete, specific version of that same observation.

Should You Upgrade?

  • Upgrade if your workload is primarily Chinese-language (the tokenizer change barely affects Chinese token counts), or if you're doing agentic/tool-heavy work where the Terminal-Bench and SWE-bench gains matter more than a modest cost increase.

  • Hold off if your workload is high-volume and English-heavy, and Sonnet 4.6 is already meeting your quality bar — re-run your own cost math rather than trusting the sticker price, especially once the introductory pricing window closes on August 31, 2026.

  • Either way, don't assume the newer, cheaper-looking model automatically costs less per task. On our test it didn't.

FAQ

Is Sonnet 5 really a "strict improvement" over Sonnet 4.6, like Anthropic says?

On the capability benchmarks Anthropic has published, yes — we didn't find one where Sonnet 4.6 leads. On cost per task, our test found the opposite once you factor in the tokenizer change, so "strict improvement" doesn't extend to cost-efficiency for every workload.

Why does Sonnet 5 feel "more on edge" than Sonnet 4.6?

Anthropic hasn't published specifics on this. From our own test outputs, Sonnet 5 was more willing to add unrequested scope (thread-safety, a more defensive test-clock setup) than Sonnet 4.6, which stayed closer to the literal request. That's consistent with — though narrower than — the community description.

Is Sonnet 5 the default model now? Can I still use Sonnet 4.6?

Per Anthropic's launch announcement, Sonnet 5 replaced Sonnet 4.6 as the default for Free and Pro users on claude.ai. Max, Team, and Enterprise users, along with API users, can still select Sonnet 4.6 directly.

Should I compare Sonnet 5 to Sonnet 4.6 or to Opus 4.8?

Different decision entirely. This page is about the generational upgrade; if you're choosing between Sonnet 5 and Opus 4.8 specifically, we ran a separate set of head-to-head tests — see Sonnet 5 vs Opus 4.8: real tests.

Is the "higher effort tier is cheaper" claim people are posting online real?

Not in our test. We ran Sonnet 5 at medium and high against Sonnet 4.6 at low and medium, and Sonnet 5 cost more at every pairing. Individual anecdotal reports vary by task — run your own comparison on your actual workload before assuming a specific effort-level crossover.

Appendix: Exact Prompt and Raw Output

The task we sent to both models:

Implement a token-bucket rate limiter as a Python class `RateLimiter(capacity: int,
refill_rate: float)` with a method `allow() -> bool` that returns whether a request
is allowed right now, consuming one token if so. Use a monotonic clock, no external
dependencies. Save it to limiter.py.

Then write test_limiter.py with at least 5 test cases covering: burst up to capacity
succeeds then blocks, tokens refill over time (use time.sleep or a mockable clock),
refill rate is respected (not instant full refill), capacity is never exceeded, and
zero/negative capacity is handled sanely.

Run the tests with pytest and make sure they all pass. If any fail, fix the code and
re-run until all pass. Report the final pytest output.

The actual API response for the Sonnet 5 (medium) run, trimmed to the fields relevant to cost and timing:

{
  "duration_ms": 31616,
  "num_turns": 5,
  "stop_reason": "end_turn",
  "total_cost_usd": 0.3438645,
  "usage": {
    "input_tokens": 4302,
    "cache_creation_input_tokens": 60894,
    "cache_read_input_tokens": 233420,
    "output_tokens": 2172
  }
}