On August 10, 2026, Anthropic cancelled Claude Sonnet 5's scheduled September 1 price increase - $2 per million input tokens and $10 per million output tokens is now the permanent standard rate, not an introductory price. If your Q4 API budget was built on the $3/$15 figures from the June launch coverage, one line changes; and if you stayed on Sonnet 4.6 waiting out the "real" price, the answer just changed too.
Claude Sonnet 5 pricing right now: the full rate card
Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens on Anthropic's first-party API, and every discounted rate scales from that base. The Claude Platform pricing documentation, checked August 16, 2026, states it plainly: "The previously scheduled increase ... will not occur."
| Rate | Price per 1M tokens | Multiplier vs. base input |
|---|---|---|
| Base input | $2.00 | 1.0x |
| Base output | $10.00 | 5x input |
| 5-min cache write | $2.50 | 1.25x |
| 1-hour cache write | $4.00 | 2x |
| Cache read (hit) | $0.20 | 0.1x |
| Batch input | $1.00 | 50% off |
| Batch output | $5.00 | 50% off |
Three details from the same docs that don't fit in a headline rate. The full 1M-token context window bills at the same per-token price - a 900k-token request and a 9k-token request cost the same per token, with no long-context surcharge. Pinning inference to the US with inference_geo: "us" adds a 1.1x multiplier across input, output, and cache rates. And server-side web search bills at $10 per 1,000 searches on top of normal tokens.
The September 1 price increase that never happened
Claude Sonnet 5 launched June 30, 2026 with $2/$10 explicitly labeled introductory pricing, set to revert to $3/$15 - the same rate as Sonnet 4.6 - on September 1, 2026. That reversion is the number most pricing tables published in early July still carry. On August 10, 2026, the launch announcement's changelog was updated with a single line: introductory pricing became permanent, and TechCrunch's launch piece now carries an update note reflecting the standing rate.
Treat any table that still describes the $3/$15 step-up as upcoming or "after August 31" as pre-August 10 - it overstates both standard rates by 50%. Beyond the deadline, nothing structural changed: Sonnet 5 at $2/$10 sits one tier below Sonnet 4.6's $3/$15, a 33% input cut.
Sonnet 5 vs Sonnet 4.6: the migration math flipped
The original case for caution was the tokenizer. Sonnet 5 ships with the updated tokenizer family introduced with Opus 4.7 - Anthropic's launch announcement puts the token multiplier for identical text at 1.0x-1.35x depending on content type, and the platform docs put the typical increase near 30%. Under the old plan, $3/$15 standard pricing would have made an equivalent prompt cost up to $4.05 per million tokens of input and $20.25 of output against Sonnet 4.6's $3/$15.
At the permanent $2/$10, the worst case inverts the comparison. Even at maximum 1.35x token inflation, equivalent text costs $2.70 per million input tokens and $13.50 per million output tokens - below Sonnet 4.6 on both axes, before counting that Sonnet 5 scores higher on the launch comparison table: 63.2% vs 58.1% on SWE-bench Pro, 80.4% vs 67.0% on Terminal-Bench 2.1, and 81.2% vs 78.5% on OSWorld-Verified.
The counterargument worth keeping on your radar is cost per solved problem, not per token. Artificial Analysis testing cited by eesel AI put Sonnet 5 at roughly $2.29 per task on its Intelligence Index and found that without promotional pricing it could cost more per task than Opus 4.8, because high-effort runs burn thinking tokens billed at the output rate. The community reaction was blunt:
"tl;dr: Sonnet 5 is cheaper per token, but more expensive per solved problem - and still lags behind Opus 4.8 in overall intelligence." - u/Chubby, r/accelerate
The fixed rate changes that equation's inputs, not its lesson: effort settings and tokens-per-task still decide your bill. The capability side of the migration question is covered in our Sonnet 5 vs Sonnet 4.6 comparison.
Same headline price, different bills: effort and caching
Two levers move a Sonnet 5 invoice more than any price announcement will. First, the effort dial. Sonnet 5 exposes low, medium, high, xhigh, and max effort levels; the launch announcement claims substantially improved cost efficiency at medium effort, which launch coverage maps as Sonnet 5 at medium roughly equaling Sonnet 4.6 at high - which suggests testing whether one effort step lower preserves required quality on your workload. Second, prompt caching. A 5-minute cache write at $2.50 pays for itself after one subsequent read; a 1-hour write at $4.00 after two; every cache read then costs $0.20, a 90% discount on input. Per the platform docs, cache and batch multipliers stack.
Batch pricing deserves a closer look on a permanent base. At $1/$5 per million tokens, asynchronous Sonnet 5 processing (results within 24 hours) costs the same as Haiku 4.5's standard on-demand rate - for a model that clears 63.2% on SWE-bench Pro. Batch is usually the cheapest route for independent, non-latency-sensitive evaluation runs, backfills, and document processing.
Same model, different bills across providers
Sonnet 5's rate card is not uniform across the seven providers that host it. CloudPrice's provider table shows Anthropic direct, Amazon Bedrock, Azure AI Foundry, Google Vertex AI, OpenRouter, and Vercel AI Gateway all at $2.00/$10.00 - while Snowflake charges $2.40/$12.00, a 20% premium. Two adders apply regardless of vendor, per the platform docs: regional or multi-region endpoints on Bedrock and Vertex cost about 10% more than global endpoints, and US-only inference adds the 1.1x multiplier.
| Provider | Input / 1M | Output / 1M | Notes |
|---|---|---|---|
| Anthropic API | $2.00 | $10.00 | Batch $1/$5, full cache rates |
| Amazon Bedrock | $2.00 | $10.00 | Regional endpoints +10% |
| Google Vertex AI | $2.00 | $10.00 | Regional endpoints +10% |
| Azure AI Foundry | $2.00 | $10.00 | US Data Zone deployment available |
| OpenRouter | $2.00 | $10.00 | Batch variant $1/$5 |
| Vercel AI Gateway | $2.00 | $10.00 | - |
| Snowflake | $2.40 | $12.00 | +20% vs. standard rate |
One trap when comparing across directories: the "starting at $1.00/$5.00" figure some listings attach to Sonnet 5 is the batch route, not the on-demand price - the same directory lists Sonnet 4.6 at $1.50/$7.50, which is exactly Sonnet 4.6's batch rate rather than its $3/$15 standard.
What permanent $2/$10 means for your budget
The decisions that remain are model-level, not deadline-driven. Sonnet 5 is the default model on Claude Free and Pro plans and is included in Max, Team, and Enterprise, per the launch announcement - subscription users never touch these API rates, and the Claude API pricing guide covers where the subscription-to-API crossover lands. For API buyers:
| Situation | Call |
|---|---|
| Production agents, latency-sensitive | Sonnet 5 at $2/$10, 5-min cache on system prompts |
| Overnight backfills, eval runs, bulk processing | Sonnet 5 Batch at $1/$5 |
| The hardest coding tasks (Opus 4.8: 69.2% vs 63.2% on SWE-bench Pro, per Anthropic's launch comparison table) | Opus 4.8 at $5/$25 |
| Classification, routing, summarization | Haiku 4.5 at $1/$5 |
| Still running Sonnet 4.6 in production | Pilot the migration - Anthropic's documented 1.0-1.35x tokenizer range still lands below 4.6 per equivalent token; validate quality and tokens per task on real traffic first |
At xhigh and max, thinking tokens billed at $10 per million still scale the bill with problem difficulty. Artificial Analysis found a high-effort Sonnet 5 run could cost more per task than Opus 4.8 at the old $3/$15 rates; the permanent cut lowers that floor by a third but does not remove it - measure tokens per task before committing a quarterly number. Run a week of real traffic through the Claude API endpoint at your production effort setting and price the workload, not just dollars per million tokens.
Related reading: Claude Sonnet 5 vs Sonnet 4.6 - Claude API Pricing Guide 2026 - Claude Sonnet 5 Guide