AIREITER
AnthropicText Chat

Claude Sonnet 5.5

Claude Sonnet 5.5 API at half of Anthropic list: $1 in, $5 out per 1M tokens. What changed from Sonnet 5, benchmarks, breaking changes, and working code.

InputOfficial $2.00 per 1M tokensAIReiter $1.00 per 1M tokensOutputOfficial $10.00 per 1M tokensAIReiter $5.00 per 1M tokensCache readOfficial $0.20 per 1M tokensAIReiter $0.10 per 1M tokensCache creationOfficial $2.50 per 1M tokensAIReiter $1.25 per 1M tokens
Run with API

INPUT

OUTPUT

Example
Generated in
42.7 seconds
Input tokens
134
Output tokens
2354
Tokens per second
55.13 tokens / second
Time to first token
-

Model details

Use the same model key in Playground, API requests, and internal workflows.

Model ID
claude-sonnet-5-5
Provider
Anthropic
Protocol
Anthropic Messages
Context window
1,000,000 tokens
Max output
128,000 tokens
Input tokens
100 credits / 1M tokens
Output tokens
500 credits / 1M tokens
Cache read
10 credits / 1M tokens
Cache write
125 credits / 1M tokens

What's new in Claude Sonnet 5.5

Anthropic released Claude Sonnet 5.5 on September 28, 2026 as a direct upgrade to Sonnet 5. The list price did not move; the work it can do, and how many tokens it spends doing it, did.

Close to Opus 5.5 on knowledge work

Anthropic reports a GDPval-AA score of 1844 for Sonnet 5.5, two points behind Opus 5.5 (1846) and about 400 points ahead of Sonnet 5 (1449). Anthropic also says Opus 5.5 is still clearly stronger on open-ended work that needs sustained judgment.

A large step up in coding

On CursorBench 4.0 Anthropic reports 55.5% for Sonnet 5.5, against 34.1% for Sonnet 5 and 57.8% for Opus 5.5. Anthropic positions it for well-scoped everyday work: building features, fixing bugs, and producing documents, slides and spreadsheets.

Fewer tokens for the same job

Early customers reported the savings in their own workloads: Slack saw about 14% fewer output tokens with no prompt changes, and Balyasny Asset Management measured roughly 121K tokens per answer against 497K on Sonnet 5 across 2,441 finance tasks.

Faster output, same tokenizer

Anthropic says output is more than 30% faster than Sonnet 5; Artificial Analysis measured about 139 output tokens per second. The tokenizer is unchanged, so the same text counts as the same number of tokens on both models.

Claude Sonnet 5.5 vs Sonnet 5 vs Opus 5.5

All three run on the same AIReiter key. The short version: move off Sonnet 5 once your requests pass the migration checks below, and keep Opus 5.5 for the hardest open-ended work.

Sonnet 5.5 vs Sonnet 5: same price, more done

Both list at $2 input and $10 output per million tokens. Sonnet 5.5 scores higher on the coding and knowledge-work benchmarks above and usually finishes in fewer tokens, so the same job tends to cost less.

Sonnet 5.5 vs Opus 5.5: half the price

Here Sonnet 5.5 is $1.00 in and $5.00 out per million tokens; Opus 5.5 is $2.00 and $10.00. Both have a 1M-token context window and 128K max output, so switching is a change to the model field and nothing else.

When Opus 5.5 is still worth it

Long, ambiguous tasks where the model has to plan, notice its own mistakes and keep going. Base44 reported 3.6 iterations per app build on Sonnet 5.5 against 7.7 on Opus 5, so try Sonnet 5.5 first and escalate only the tasks that fail.

When to stay on Sonnet 5 for now

If your code disables thinking, forces a specific tool, or sets temperature, those requests will fail on Sonnet 5.5 until you change them. Security tooling may also see more refusals. Fix those first, then switch.

Migrating from Sonnet 5: what now returns an error

Changing the model string to claude-sonnet-5-5 is not always enough. These request shapes worked on Sonnet 5 and are rejected by Sonnet 5.5.

thinking: disabled returns 400

Adaptive thinking is always on. Replace {"type": "disabled"} with {"type": "between_tools"}, the lowest setting, which skips thinking before the first response. between_tools only works at low, medium or high effort; for xhigh or max, omit thinking or send {"type": "adaptive"}.

Manual thinking budgets return 400

thinking: {"type": "enabled", "budget_tokens": N} is rejected. Control depth with the effort level instead. Thinking tokens still come out of max_tokens and bill as output, so raise max_tokens if answers start getting cut off.

Forced tool choice is rejected

A tool_choice that forces one specific tool now errors. Use automatic tool choice with strict tool definitions, and let the model decide when to call.

Custom temperature, top_p and top_k return 400

Only default sampling is accepted. Remove these fields from the request rather than setting them to fixed values. One change works in your favour: the minimum cacheable prompt drops to 512 tokens from 1,024.

Pick the effort level before you ship

Effort is now the main dial for cost and latency. There are five levels: low, medium, high, xhigh and max.

The API default is high

If you send no effort, the API runs at high. Claude Code and the Claude apps default to medium. Set output_config.effort explicitly so your bill does not depend on a default.

max is expensive

At max effort Artificial Analysis measured $7.60 per task at list price and 410M output tokens across its index, against a median of 88M. Anthropic's efficiency claims are about the lower levels, not max.

Start at medium, measure, then move

Run your real tasks at medium, check quality, and step up only where it falls short. For latency-sensitive tool loops, low or medium with between_tools thinking keeps first responses quick.

Claude Sonnet 5.5 API pricing on AIReiter

Every line is billed at 50% of Anthropic list. Same model, same Messages protocol.

Half of list on every line

Input $1.00 (list $2.00). Output $5.00 (list $10.00). Cache reads $0.10 (list $0.20). Cache writes $1.25 (list $2.50). All per million tokens.

A real call: $1.50 instead of $3.00

One million input tokens plus one hundred thousand output tokens costs $1.50 here and $3.00 at Anthropic list. The same call on Opus 5.5 here costs $3.00.

Billed on tokens actually used

An estimate is held when the call starts, then settled against the real token count when the response completes, and the difference is returned. The rate is the same from the first token to the millionth; there is no long-context tier.

How to call the Claude Sonnet 5.5 API

The endpoint speaks the Anthropic Messages protocol, so an existing Claude client needs only a new base URL and key.

01

Get a key and point the client at AIReiter

Create an API key in your AIReiter dashboard. In the official Anthropic SDKs set the base URL to https://aireiter.com/api — the SDK adds /v1/messages itself.

02

Send a request with curl

curl https://aireiter.com/api/v1/messages \ -H "x-api-key: $AIREITER_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-5-5", "max_tokens": 4096, "output_config": {"effort": "medium"}, "messages": [{"role": "user", "content": "Hello"}] }'

03

Or use the Python SDK

from anthropic import Anthropic client = Anthropic(api_key="YOUR_AIREITER_KEY", base_url="https://aireiter.com/api") msg = client.messages.create( model="claude-sonnet-5-5", max_tokens=4096, messages=[{"role": "user", "content": "Hello"}], ) print(msg.content[0].text)

Claude Sonnet 5.5 API FAQ

/ 01

When was Claude Sonnet 5.5 released, and what is the model ID?

Anthropic released it on September 28, 2026. The API model ID is claude-sonnet-5-5, and its knowledge cutoff is June 2026.

/ 02

What is the context window and max output?

A 1M-token context window and up to 128K output tokens per request. It accepts text and image input and returns text.

/ 03

How much does the Claude Sonnet 5.5 API cost?

Anthropic lists $2 input and $10 output per million tokens, the same as Sonnet 5. AIReiter charges half: $1 input, $5 output, $0.10 cache reads and $1.25 cache writes.

/ 04

Can I turn thinking off?

Not fully. thinking disabled returns a 400 error. The lowest setting is between_tools, which works at low, medium and high effort.

/ 05

Is Sonnet 5.5 better than Sonnet 5?

On Anthropic's reported coding and knowledge-work benchmarks, yes, by a wide margin, at the same list price. Check the migration changes above before switching, because some Sonnet 5 request shapes now fail.

/ 06

Why was my security question refused or answered by Sonnet 5?

Sonnet 5.5 is the first Sonnet to ship with the cybersecurity safeguards used on Anthropic's most capable models. Some higher-risk requests fall back to Sonnet 5, and benign security tasks can see more refusals. This is upstream behaviour.

/ 07

Do I need an Anthropic account?

No. An AIReiter key works against the AIReiter endpoint, and you can try the model in the playground on this page before writing any code.

/ 08

Which APIs are supported?

Anthropic Messages at /api/v1/messages is the primary contract, with streaming. OpenAI-style Chat Completions and the Responses API are also served, and /api/v1/models lists the models available to your key.