Claude Sonnet 5.5 combines $2-per-million input pricing with a published 70.6% Terminal-Bench 4.0 score, but high-effort agent runs can make its real cost per accepted change much higher. For teams choosing a production coding model, Sonnet 5.5 is the default lane to pilot, not an automatic replacement for Opus 5.5.
The production verdict: use Sonnet 5.5 as the default coding lane
Pilot Claude Sonnet 5.5 first for scoped bug fixes, refactors, test generation, and tool-using repository tasks. Anthropic’s published Terminal-Bench 4.0 result is 70.6%, above Claude Opus 5.5’s 66.4% and Claude Sonnet 5’s 10.3% in the same comparison.
Keep one guardrail: set effort and output limits explicitly. Independent cost analyses report that Sonnet 5.5 at max effort can cost more per benchmark task than Opus 5.5, despite Sonnet’s lower headline rates.
What Claude Sonnet 5.5 changes for API teams
Claude Sonnet 5.5 launched on September 28, 2026. The official model documentation lists model ID claude-sonnet-5-5, a 1-million-token context window, 128,000-token standard maximum output, adaptive thinking, and a default API effort level of high.
| Production detail | Claude Sonnet 5.5 |
|---|---|
| Release date | September 28, 2026 |
| Model ID | claude-sonnet-5-5 |
| Context window | 1M tokens |
| Standard maximum output | 128K tokens |
| Batch maximum output | 300K tokens with the documented beta header |
| API default effort | high |
Anthropic’s documentation commits not to retire the model before September 28, 2027. That is a lifecycle floor, not a promised final retirement date.
The API defaults that affect a coding bill
Adaptive thinking is enabled by default. Teams migrating from Sonnet 5 should test between_tools if they need to disable up-front thinking; the official documentation says forced tool use now returns an error, and non-default temperature, top_p, or top_k values return HTTP 400 errors.
Text generated between tool calls may arrive in thinking blocks. A streaming client that assumes every intermediate message is a normal text block can appear silent after migration. Update the parser before routing Sonnet 5.5 into an existing coding agent.
Terminal-Bench performance: strong enough for the default lane
The most prominent coding-agent benchmark cited here is Terminal-Bench 4.0, which evaluates multi-step command-line tasks. Anthropic-reported results put Sonnet 5.5 at 70.6%, compared with 66.4% for Opus 5.5 and 10.3% for Sonnet 5.
| Model | Terminal-Bench 4.0 | GDPval-AA Elo | CursorBench 4.0 |
|---|---|---|---|
| Claude Sonnet 5.5 | 70.6% | 1844 | 55.5% |
| Claude Opus 5.5 | 66.4% | 1846 | 57.8% |
| Claude Sonnet 5 | 10.3% | 1449 | 34.1% |
The benchmark figures above are reported in the DataCamp benchmark summary, which attributes the evaluation results to Anthropic’s launch material. Sonnet 5.5 leads the terminal-coding comparison, but Opus 5.5 remains ahead on CursorBench and several broader reasoning and knowledge-work evaluations.
Validate on repository replays using passing tests, tool turns, and accepted changes, not diffs alone.
Claude Sonnet 5.5 API pricing in real workloads
The Anthropic rate card is straightforward, but a coding agent may pay for more than the visible prompt. Thinking tokens are billed as output, and repeated repository context can become cache reads across turns.
| API item | Claude Sonnet 5.5 price |
|---|---|
| Input | $2 per 1M tokens |
| Output, including thinking | $10 per 1M tokens |
| 5-minute cache write | $2.50 per 1M tokens |
| 1-hour cache write | $4 per 1M tokens |
| Cache read | $0.20 per 1M tokens |
| Batch input | 50% discount, equivalent to $1 per 1M |
| Batch output | 50% discount, equivalent to $5 per 1M |
The official docs list a 512-token minimum cacheable prompt and explain that Sonnet 5.5 uses adaptive thinking. Sonnet 5.5 uses the same tokenizer as Sonnet 5, according to the independent pricing analysis from eesel; a migration does not automatically reduce token counts.
A request containing 4,000 input tokens and 700 output tokens costs about $0.015 before other charges: $0.008 for input plus $0.007 for output. A coding agent that takes 20 turns with 3,000 fresh input tokens and 2,000 output tokens per turn would consume roughly $0.12 in fresh input and $0.40 in output before cache reads, cache writes, tools, or retries. These are workload illustrations, not universal per-task prices.
Effort is the price control
The TokenCost effort analysis, based on Artificial Analysis Intelligence Index measurements, reported these Sonnet 5.5 results:
| Effort | Score | Full-index cost |
|---|---|---|
| Low | 35.8 | $544 |
| Medium | 40.7 | $701 |
| High | 46.7 | $1,176 |
| Xhigh | 51.9 | $2,738 |
| Max | 56.0 | $8,977 |
The max result is the warning. The same analysis put Opus 5.5 max at 57.6 for $8,708, while Opus 5.5 xhigh reached 56.0 for $4,057. Sonnet 5.5 max used approximately 193,000 output tokens per task in that test.
Start at medium or high, cap output tokens, and promote only failed or high-risk tasks to a more expensive route. Do not copy an old Sonnet 5 max setting into Sonnet 5.5 without rerunning cost and quality evaluations.
Migration checklist for a production coding model
- Pin
claude-sonnet-5-5in staging rather than changing an alias globally. - Replay representative bug fixes, refactors, tests, and multi-file changes from your repositories.
- Set effort explicitly and record output tokens, cache reads, tool turns, wall-clock time, and accepted-change rate.
- Replace
thinking: disabledwith the supportedbetween_toolsbehavior where appropriate. - Remove forced tool-use assumptions and test the new tool-choice behavior.
- Update streaming code to handle
thinkingblocks between tool calls. - Recheck non-default sampling parameters; the official docs list 400 errors for non-default
temperature,top_p, andtop_k. - Add a spend ceiling and an abort condition for runaway output or repeated tool loops.
- Compare cost per accepted change, not cost per request.
- Roll out to a small percentage of traffic. Promote Sonnet only if it matches the incumbent’s accepted-change rate within your agreed tolerance while reducing cost per accepted change.
Anthropic employee @cjav_dev reported that requests using thinking: {"type":"disabled"} began returning 400 errors and should move to between_tools instead (post on X).
When to choose Sonnet 5.5, Opus 5.5, or a cheaper model
| Workload | Recommended first choice | Why |
|---|---|---|
| Scoped bug fixes and refactors | Sonnet 5.5 at medium/high | Strong terminal score with lower token rates |
| High-volume code classification or simple edits | Sonnet 5.5 at low/medium, or a cheaper model | Avoid paying for unnecessary reasoning |
| Long-running repository agent | Sonnet 5.5 with caching and hard budgets | Cache and turn counts decide the real bill |
| Ambiguous architecture or final review | Opus 5.5 | Broader judgment matters more than the lowest rate card |
| Offline, non-urgent code analysis | Sonnet 5.5 Batch API | The API offers a 50% input/output discount |
| Max-effort terminal experiment | Sonnet 5.5 only after an internal benchmark | Terminal-Bench is a relative strength, but max can be expensive |
Use Sonnet for well-defined, measurable tasks; escalate ambiguous architectural work to Opus.
Claude Sonnet 5.5 API FAQ
What is Claude Sonnet 5.5 API pricing?
The standard rate is $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million tokens, cache writes cost $2.50 for five minutes or $4 for one hour, and the Batch API discounts input and output by 50%.
Is Sonnet 5.5 better than Opus 5.5 for coding?
Sonnet 5.5 leads the published Terminal-Bench 4.0 comparison at 70.6% versus Opus 5.5 at 66.4%. Opus remains ahead on several other evaluations, so teams should route by task type rather than treat the coding result as a universal ranking.
What is the Sonnet 5.5 model ID?
Use claude-sonnet-5-5 on the Claude API. Provider-specific identifiers are listed in Anthropic’s model documentation.
Does Sonnet 5.5 support a 1M-token context window?
Yes. Anthropic lists a 1-million-token context window and a 128,000-token standard maximum output. The Message Batches API beta can support 300,000 output tokens with the documented beta header.
What changed when migrating from Sonnet 5?
Adaptive thinking and response-block behavior changed, forced tool use can error, non-default sampling parameters can return 400 errors, and thinking: disabled must be replaced with the supported behavior. Re-run agent and streaming tests before production rollout.
Run a one-week medium/high-effort replay pilot, measuring cost per accepted change and escalating failed cases to Opus.