AIREITER

Claude Sonnet 5.5 API Review and Pricing for Coding Teams

Last Updated: 2026-09-30 00:38:01

Claude Sonnet 5.5 combines $2-per-million input pricing with a published 70.6% Terminal-Bench 4.0 score, but high-effort agent runs can make its real cost per accepted change much higher. For teams choosing a production coding model, Sonnet 5.5 is the default lane to pilot, not an automatic replacement for Opus 5.5.

The production verdict: use Sonnet 5.5 as the default coding lane

Pilot Claude Sonnet 5.5 first for scoped bug fixes, refactors, test generation, and tool-using repository tasks. Anthropic’s published Terminal-Bench 4.0 result is 70.6%, above Claude Opus 5.5’s 66.4% and Claude Sonnet 5’s 10.3% in the same comparison.

Keep one guardrail: set effort and output limits explicitly. Independent cost analyses report that Sonnet 5.5 at max effort can cost more per benchmark task than Opus 5.5, despite Sonnet’s lower headline rates.

What Claude Sonnet 5.5 changes for API teams

Claude Sonnet 5.5 launched on September 28, 2026. The official model documentation lists model ID claude-sonnet-5-5, a 1-million-token context window, 128,000-token standard maximum output, adaptive thinking, and a default API effort level of high.

Production detailClaude Sonnet 5.5
Release dateSeptember 28, 2026
Model IDclaude-sonnet-5-5
Context window1M tokens
Standard maximum output128K tokens
Batch maximum output300K tokens with the documented beta header
API default efforthigh

Anthropic’s documentation commits not to retire the model before September 28, 2027. That is a lifecycle floor, not a promised final retirement date.

The API defaults that affect a coding bill

Adaptive thinking is enabled by default. Teams migrating from Sonnet 5 should test between_tools if they need to disable up-front thinking; the official documentation says forced tool use now returns an error, and non-default temperature, top_p, or top_k values return HTTP 400 errors.

Text generated between tool calls may arrive in thinking blocks. A streaming client that assumes every intermediate message is a normal text block can appear silent after migration. Update the parser before routing Sonnet 5.5 into an existing coding agent.

Terminal-Bench performance: strong enough for the default lane

The most prominent coding-agent benchmark cited here is Terminal-Bench 4.0, which evaluates multi-step command-line tasks. Anthropic-reported results put Sonnet 5.5 at 70.6%, compared with 66.4% for Opus 5.5 and 10.3% for Sonnet 5.

ModelTerminal-Bench 4.0GDPval-AA EloCursorBench 4.0
Claude Sonnet 5.570.6%184455.5%
Claude Opus 5.566.4%184657.8%
Claude Sonnet 510.3%144934.1%

The benchmark figures above are reported in the DataCamp benchmark summary, which attributes the evaluation results to Anthropic’s launch material. Sonnet 5.5 leads the terminal-coding comparison, but Opus 5.5 remains ahead on CursorBench and several broader reasoning and knowledge-work evaluations.

Claude Sonnet 5.5 coding benchmark and index cost comparison

Validate on repository replays using passing tests, tool turns, and accepted changes, not diffs alone.

Claude Sonnet 5.5 API pricing in real workloads

The Anthropic rate card is straightforward, but a coding agent may pay for more than the visible prompt. Thinking tokens are billed as output, and repeated repository context can become cache reads across turns.

API itemClaude Sonnet 5.5 price
Input$2 per 1M tokens
Output, including thinking$10 per 1M tokens
5-minute cache write$2.50 per 1M tokens
1-hour cache write$4 per 1M tokens
Cache read$0.20 per 1M tokens
Batch input50% discount, equivalent to $1 per 1M
Batch output50% discount, equivalent to $5 per 1M

The official docs list a 512-token minimum cacheable prompt and explain that Sonnet 5.5 uses adaptive thinking. Sonnet 5.5 uses the same tokenizer as Sonnet 5, according to the independent pricing analysis from eesel; a migration does not automatically reduce token counts.

A request containing 4,000 input tokens and 700 output tokens costs about $0.015 before other charges: $0.008 for input plus $0.007 for output. A coding agent that takes 20 turns with 3,000 fresh input tokens and 2,000 output tokens per turn would consume roughly $0.12 in fresh input and $0.40 in output before cache reads, cache writes, tools, or retries. These are workload illustrations, not universal per-task prices.

Effort is the price control

The TokenCost effort analysis, based on Artificial Analysis Intelligence Index measurements, reported these Sonnet 5.5 results:

EffortScoreFull-index cost
Low35.8$544
Medium40.7$701
High46.7$1,176
Xhigh51.9$2,738
Max56.0$8,977

The max result is the warning. The same analysis put Opus 5.5 max at 57.6 for $8,708, while Opus 5.5 xhigh reached 56.0 for $4,057. Sonnet 5.5 max used approximately 193,000 output tokens per task in that test.

Start at medium or high, cap output tokens, and promote only failed or high-risk tasks to a more expensive route. Do not copy an old Sonnet 5 max setting into Sonnet 5.5 without rerunning cost and quality evaluations.

Migration checklist for a production coding model

  1. Pin claude-sonnet-5-5 in staging rather than changing an alias globally.
  2. Replay representative bug fixes, refactors, tests, and multi-file changes from your repositories.
  3. Set effort explicitly and record output tokens, cache reads, tool turns, wall-clock time, and accepted-change rate.
  4. Replace thinking: disabled with the supported between_tools behavior where appropriate.
  5. Remove forced tool-use assumptions and test the new tool-choice behavior.
  6. Update streaming code to handle thinking blocks between tool calls.
  7. Recheck non-default sampling parameters; the official docs list 400 errors for non-default temperature, top_p, and top_k.
  8. Add a spend ceiling and an abort condition for runaway output or repeated tool loops.
  9. Compare cost per accepted change, not cost per request.
  10. Roll out to a small percentage of traffic. Promote Sonnet only if it matches the incumbent’s accepted-change rate within your agreed tolerance while reducing cost per accepted change.

Anthropic employee @cjav_dev reported that requests using thinking: {"type":"disabled"} began returning 400 errors and should move to between_tools instead (post on X).

When to choose Sonnet 5.5, Opus 5.5, or a cheaper model

WorkloadRecommended first choiceWhy
Scoped bug fixes and refactorsSonnet 5.5 at medium/highStrong terminal score with lower token rates
High-volume code classification or simple editsSonnet 5.5 at low/medium, or a cheaper modelAvoid paying for unnecessary reasoning
Long-running repository agentSonnet 5.5 with caching and hard budgetsCache and turn counts decide the real bill
Ambiguous architecture or final reviewOpus 5.5Broader judgment matters more than the lowest rate card
Offline, non-urgent code analysisSonnet 5.5 Batch APIThe API offers a 50% input/output discount
Max-effort terminal experimentSonnet 5.5 only after an internal benchmarkTerminal-Bench is a relative strength, but max can be expensive

Use Sonnet for well-defined, measurable tasks; escalate ambiguous architectural work to Opus.

Claude Sonnet 5.5 API FAQ

What is Claude Sonnet 5.5 API pricing?

The standard rate is $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million tokens, cache writes cost $2.50 for five minutes or $4 for one hour, and the Batch API discounts input and output by 50%.

Is Sonnet 5.5 better than Opus 5.5 for coding?

Sonnet 5.5 leads the published Terminal-Bench 4.0 comparison at 70.6% versus Opus 5.5 at 66.4%. Opus remains ahead on several other evaluations, so teams should route by task type rather than treat the coding result as a universal ranking.

What is the Sonnet 5.5 model ID?

Use claude-sonnet-5-5 on the Claude API. Provider-specific identifiers are listed in Anthropic’s model documentation.

Does Sonnet 5.5 support a 1M-token context window?

Yes. Anthropic lists a 1-million-token context window and a 128,000-token standard maximum output. The Message Batches API beta can support 300,000 output tokens with the documented beta header.

What changed when migrating from Sonnet 5?

Adaptive thinking and response-block behavior changed, forced tool use can error, non-default sampling parameters can return 400 errors, and thinking: disabled must be replaced with the supported behavior. Re-run agent and streaming tests before production rollout.

Run a one-week medium/high-effort replay pilot, measuring cost per accepted change and escalating failed cases to Opus.