AIREITER
AnthropicText Chat

Claude Haiku 5.5

Claude Haiku 5.5 API at half of Anthropic list: $0.05 in, $0.25 out per 1M tokens up to 100K input. Long-context tier, breaking changes from Haiku 4.5, working code.

Input tokensInputOutputCache readCache creation
≤ 100,000$0.05 per 1M tokens$0.25 per 1M tokens$0.005 per 1M tokens$0.0625 per 1M tokens
> 100,000$0.25 per 1M tokens$1.25 per 1M tokens$0.025 per 1M tokens$0.3125 per 1M tokens

Rates are per million tokens. The tier applies to the entire request based on total input tokens, including cached reads and writes.

Run with API

INPUT

OUTPUT

Example
Generated in
42.7 seconds
Input tokens
134
Output tokens
2354
Tokens per second
55.13 tokens / second
Time to first token
-

Model details

Use the same model key in Playground, API requests, and internal workflows.

Model ID
claude-haiku-5-5
Provider
Anthropic
Protocol
Anthropic Messages
Context window
1,000,000 tokens
Max output
128,000 tokens

What's new in Claude Haiku 5.5

Anthropic released Claude Haiku 5.5 on October 7, 2026, a year after Haiku 4.5. The price dropped by 90% on short prompts and the model moved much closer to Sonnet.

The first Haiku with effort control

Haiku 5.5 uses adaptive thinking and five effort levels: low, medium, high, xhigh and max. The default is medium. Earlier Haiku models had no effort setting.

1M context and 128K output

The context window grows from 200K tokens on Haiku 4.5 to 1M, and max output from 64K to 128K. It accepts text and image input. Knowledge cutoff is June 2026.

A big jump on agent and office work

At max effort Anthropic reports 72.4% on OSWorld 2.1 (Haiku 4.5: 15.7%) and a GDPval-AA Elo of 1620 (Haiku 4.5: 735). Sonnet 5.5 still leads on every benchmark Anthropic published.

Independent score

Artificial Analysis gives Haiku 5.5 a score of 43 on its Intelligence Index at max effort, up from 17 for Haiku 4.5, and 34 at the default medium effort.

Claude Haiku 5.5 vs Haiku 4.5 vs Sonnet 5.5

All of these run on the same AIReiter key, so switching is a change to the model field.

Haiku 5.5 vs Haiku 4.5

Haiku 4.5 lists at $1 input and $5 output per million tokens. Haiku 5.5 lists at $0.10 and $0.50 for prompts up to 100K tokens. Anthropic estimates real workloads cost about 75% less after counting the newer tokenizer, which produces roughly 30% more tokens for the same text.

Haiku 5.5 vs Sonnet 5.5

Sonnet 5.5 is $1.00 in and $5.00 out per million tokens here, twenty times Haiku 5.5's base rate. Anthropic still recommends Sonnet 5.5 or Opus 5.5 for complex coding. Use Haiku 5.5 for classification, extraction, summaries, routing and subagent steps.

Haiku 5.5 vs GPT-6 Luna

Both list at $0.10 input and $0.50 output. Anthropic reports Haiku 5.5 ahead on OSWorld 2.1 (72.4% vs 48.9%) and GDPval-AA (1620 vs 1437). Artificial Analysis notes Haiku 5.5 uses more output tokens than Luna to reach similar scores, so compare on your own tasks.

Claude Haiku 5.5 API pricing on AIReiter

Every line is billed at 50% of Anthropic list. The rate depends on how many input tokens a single request sends.

Up to 100K input tokens

Input $0.05 (list $0.10). Output $0.25 (list $0.50). Cache reads $0.005 (list $0.01). Cache writes $0.0625 (list $0.125). All per million tokens.

Over 100K input tokens: every rate x5

Input $0.25 (list $0.50). Output $1.25 (list $2.50). Cache reads $0.025 (list $0.05). Cache writes $0.3125 (list $0.625). The tier is decided per request by total input, including cached tokens, and applies to the whole request.

Two real calls

50K input plus 5K output costs $0.00375 here ($0.0075 at list). 200K input plus 10K output falls in the long-context tier and costs $0.0625 here ($0.125 at list). If you can split a long job into requests under 100K, it costs a fifth as much.

Billed on tokens actually used

An estimate is held when the call starts, then settled against the real token count when the response completes, and the difference is returned. Thinking tokens bill as output.

Migrating from Haiku 4.5: what now returns an error

Changing the model string to claude-haiku-5-5 is not always enough. These request shapes worked on Haiku 4.5 and are rejected on Haiku 5.5.

Drop temperature, top_p and top_k

Non-default sampling values return 400 with "deprecated for this model". Remove these fields and steer output with the prompt and the effort level.

No assistant prefill

A conversation that ends with an assistant message returns 400. The last message must be from the user. To force a format, use structured outputs or say it in the prompt.

Replace thinking budgets with effort

Anthropic's API rejects thinking: {"type": "enabled", "budget_tokens": N} on this model, so remove it and set output_config.effort instead. Thinking tokens still come out of max_tokens, so raise max_tokens if answers get cut off.

Turning thinking off caps effort at high

thinking: {"type": "disabled"} works at low, medium and high effort. With xhigh or max it returns 400. The between_tools thinking mode used on Sonnet 5.5 is not supported on Haiku 5.5. Forced tool_choice still works.

Pick the effort level before you ship

Effort is the main dial for quality, latency and cost. On Haiku 5.5 the gap between levels is large.

Headline scores are at max effort

Anthropic's launch numbers use max effort. At the default medium effort Anthropic reports a GDPval-AA Elo of 1277 instead of 1620. Read benchmark claims with the effort level in mind.

max is token-hungry

Artificial Analysis measured about 162K output tokens per Intelligence Index task at max effort, roughly three times GPT-6 Luna at max. At medium it measured about 137 output tokens per second.

Start at low or medium

For classification, extraction and routing, low or medium with thinking disabled keeps latency and cost down. Move up only on the tasks where quality falls short.

How to call the Claude Haiku 5.5 API

The endpoint speaks the Anthropic Messages protocol, so an existing Claude client needs only a new base URL and key.

01

Get a key and point the client at AIReiter

Create an API key in your AIReiter dashboard. In the official Anthropic SDKs set the base URL to https://aireiter.com/api — the SDK adds /v1/messages itself.

02

Send a request with curl

curl https://aireiter.com/api/v1/messages \ -H "x-api-key: $AIREITER_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-haiku-5-5", "max_tokens": 4096, "output_config": {"effort": "low"}, "messages": [{"role": "user", "content": "Hello"}] }'

03

Or use the Python SDK

from anthropic import Anthropic client = Anthropic(api_key="YOUR_AIREITER_KEY", base_url="https://aireiter.com/api") msg = client.messages.create( model="claude-haiku-5-5", max_tokens=4096, messages=[{"role": "user", "content": "Hello"}], ) print(next(b.text for b in msg.content if b.type == "text"))

Claude Haiku 5.5 API FAQ

/ 01

When was Claude Haiku 5.5 released, and what is the model ID?

Anthropic released it on October 7, 2026. The API model ID is claude-haiku-5-5, and its knowledge cutoff is June 2026. Anthropic says it will not be retired before October 7, 2027.

/ 02

What is the context window and max output?

A 1M-token context window and up to 128K output tokens per request. It accepts text and image input and returns text.

/ 03

How much does the Claude Haiku 5.5 API cost?

Anthropic lists $0.10 input and $0.50 output per million tokens for requests up to 100K input tokens, and five times that above 100K. AIReiter charges half: $0.05 and $0.25, or $0.25 and $1.25 above 100K.

/ 04

Why did one request cost five times more?

Its total input passed 100K tokens, so the whole request billed at the long-context rate. Cached tokens count toward the 100K limit too.

/ 05

Why do my answers start with a thinking block?

Adaptive thinking is on by default, so the response can contain a thinking block before the text. Read content blocks by type instead of taking content[0], or disable thinking at low, medium or high effort.

/ 06

Is Haiku 5.5 good enough for coding?

For small, well-scoped edits and subagent steps, often yes. For complex multi-file work Anthropic recommends Sonnet 5.5 or Opus 5.5, both available on the same key.

/ 07

Do I need an Anthropic account?

No. An AIReiter key works against the AIReiter endpoint, and you can try the model in the playground on this page before writing any code.

/ 08

Which APIs are supported?

Anthropic Messages at /api/v1/messages is the primary contract, with streaming. OpenAI-style Chat Completions and the Responses API are also served, and /api/v1/models lists the models available to your key.