What's new in Claude Haiku 5.5
Anthropic released Claude Haiku 5.5 on October 7, 2026, a year after Haiku 4.5. The price dropped by 90% on short prompts and the model moved much closer to Sonnet.
The first Haiku with effort control
Haiku 5.5 uses adaptive thinking and five effort levels: low, medium, high, xhigh and max. The default is medium. Earlier Haiku models had no effort setting.
1M context and 128K output
The context window grows from 200K tokens on Haiku 4.5 to 1M, and max output from 64K to 128K. It accepts text and image input. Knowledge cutoff is June 2026.
A big jump on agent and office work
At max effort Anthropic reports 72.4% on OSWorld 2.1 (Haiku 4.5: 15.7%) and a GDPval-AA Elo of 1620 (Haiku 4.5: 735). Sonnet 5.5 still leads on every benchmark Anthropic published.
Independent score
Artificial Analysis gives Haiku 5.5 a score of 43 on its Intelligence Index at max effort, up from 17 for Haiku 4.5, and 34 at the default medium effort.
Claude Haiku 5.5 vs Haiku 4.5 vs Sonnet 5.5
All of these run on the same AIReiter key, so switching is a change to the model field.
Haiku 5.5 vs Haiku 4.5
Haiku 4.5 lists at $1 input and $5 output per million tokens. Haiku 5.5 lists at $0.10 and $0.50 for prompts up to 100K tokens. Anthropic estimates real workloads cost about 75% less after counting the newer tokenizer, which produces roughly 30% more tokens for the same text.
Haiku 5.5 vs Sonnet 5.5
Sonnet 5.5 is $1.00 in and $5.00 out per million tokens here, twenty times Haiku 5.5's base rate. Anthropic still recommends Sonnet 5.5 or Opus 5.5 for complex coding. Use Haiku 5.5 for classification, extraction, summaries, routing and subagent steps.
Haiku 5.5 vs GPT-6 Luna
Both list at $0.10 input and $0.50 output. Anthropic reports Haiku 5.5 ahead on OSWorld 2.1 (72.4% vs 48.9%) and GDPval-AA (1620 vs 1437). Artificial Analysis notes Haiku 5.5 uses more output tokens than Luna to reach similar scores, so compare on your own tasks.
Claude Haiku 5.5 API pricing on AIReiter
Every line is billed at 50% of Anthropic list. The rate depends on how many input tokens a single request sends.
Up to 100K input tokens
Input $0.05 (list $0.10). Output $0.25 (list $0.50). Cache reads $0.005 (list $0.01). Cache writes $0.0625 (list $0.125). All per million tokens.
Over 100K input tokens: every rate x5
Input $0.25 (list $0.50). Output $1.25 (list $2.50). Cache reads $0.025 (list $0.05). Cache writes $0.3125 (list $0.625). The tier is decided per request by total input, including cached tokens, and applies to the whole request.
Two real calls
50K input plus 5K output costs $0.00375 here ($0.0075 at list). 200K input plus 10K output falls in the long-context tier and costs $0.0625 here ($0.125 at list). If you can split a long job into requests under 100K, it costs a fifth as much.
Billed on tokens actually used
An estimate is held when the call starts, then settled against the real token count when the response completes, and the difference is returned. Thinking tokens bill as output.
Migrating from Haiku 4.5: what now returns an error
Changing the model string to claude-haiku-5-5 is not always enough. These request shapes worked on Haiku 4.5 and are rejected on Haiku 5.5.
Drop temperature, top_p and top_k
Non-default sampling values return 400 with "deprecated for this model". Remove these fields and steer output with the prompt and the effort level.
No assistant prefill
A conversation that ends with an assistant message returns 400. The last message must be from the user. To force a format, use structured outputs or say it in the prompt.
Replace thinking budgets with effort
Anthropic's API rejects thinking: {"type": "enabled", "budget_tokens": N} on this model, so remove it and set output_config.effort instead. Thinking tokens still come out of max_tokens, so raise max_tokens if answers get cut off.
Turning thinking off caps effort at high
thinking: {"type": "disabled"} works at low, medium and high effort. With xhigh or max it returns 400. The between_tools thinking mode used on Sonnet 5.5 is not supported on Haiku 5.5. Forced tool_choice still works.
Pick the effort level before you ship
Effort is the main dial for quality, latency and cost. On Haiku 5.5 the gap between levels is large.
Headline scores are at max effort
Anthropic's launch numbers use max effort. At the default medium effort Anthropic reports a GDPval-AA Elo of 1277 instead of 1620. Read benchmark claims with the effort level in mind.
max is token-hungry
Artificial Analysis measured about 162K output tokens per Intelligence Index task at max effort, roughly three times GPT-6 Luna at max. At medium it measured about 137 output tokens per second.
Start at low or medium
For classification, extraction and routing, low or medium with thinking disabled keeps latency and cost down. Move up only on the tasks where quality falls short.
How to call the Claude Haiku 5.5 API
The endpoint speaks the Anthropic Messages protocol, so an existing Claude client needs only a new base URL and key.
Get a key and point the client at AIReiter
Create an API key in your AIReiter dashboard. In the official Anthropic SDKs set the base URL to https://aireiter.com/api — the SDK adds /v1/messages itself.
Send a request with curl
curl https://aireiter.com/api/v1/messages \ -H "x-api-key: $AIREITER_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-haiku-5-5", "max_tokens": 4096, "output_config": {"effort": "low"}, "messages": [{"role": "user", "content": "Hello"}] }'
Or use the Python SDK
from anthropic import Anthropic client = Anthropic(api_key="YOUR_AIREITER_KEY", base_url="https://aireiter.com/api") msg = client.messages.create( model="claude-haiku-5-5", max_tokens=4096, messages=[{"role": "user", "content": "Hello"}], ) print(next(b.text for b in msg.content if b.type == "text"))
Claude Haiku 5.5 API FAQ
/ 01When was Claude Haiku 5.5 released, and what is the model ID?
Anthropic released it on October 7, 2026. The API model ID is claude-haiku-5-5, and its knowledge cutoff is June 2026. Anthropic says it will not be retired before October 7, 2027.
/ 02What is the context window and max output?
A 1M-token context window and up to 128K output tokens per request. It accepts text and image input and returns text.
/ 03How much does the Claude Haiku 5.5 API cost?
Anthropic lists $0.10 input and $0.50 output per million tokens for requests up to 100K input tokens, and five times that above 100K. AIReiter charges half: $0.05 and $0.25, or $0.25 and $1.25 above 100K.
/ 04Why did one request cost five times more?
Its total input passed 100K tokens, so the whole request billed at the long-context rate. Cached tokens count toward the 100K limit too.
/ 05Why do my answers start with a thinking block?
Adaptive thinking is on by default, so the response can contain a thinking block before the text. Read content blocks by type instead of taking content[0], or disable thinking at low, medium or high effort.
/ 06Is Haiku 5.5 good enough for coding?
For small, well-scoped edits and subagent steps, often yes. For complex multi-file work Anthropic recommends Sonnet 5.5 or Opus 5.5, both available on the same key.
/ 07Do I need an Anthropic account?
No. An AIReiter key works against the AIReiter endpoint, and you can try the model in the playground on this page before writing any code.
/ 08Which APIs are supported?
Anthropic Messages at /api/v1/messages is the primary contract, with streaming. OpenAI-style Chat Completions and the Responses API are also served, and /api/v1/models lists the models available to your key.