AIREITER

Claude Haiku 5.5 API Pricing: The 100K Token Catch

Last Updated: 2026-10-07 19:00:48

Claude Haiku 5.5 looks like a straightforward low-cost model until a request crosses 100,000 input tokens. Below that boundary, Anthropic lists $0.10 per million input tokens and $0.50 per million output tokens; above it, the same request is priced at $0.50 and $2.50. That fivefold input jump, rather than the headline launch claim, is the number to put in a production budget.

Claude Haiku 5.5 API pricing at a glance

The direct API rate card separates requests by input-prompt size (USD per million tokens). The table uses Anthropic's Claude platform pricing and Haiku 5.5 launch announcement; marketplace rates can differ.

The 100K threshold is the pricing decision

The threshold concerns the request's input prompt. It is not a general surcharge applied because a model has a large context window. A request with 90,000 input tokens and 2,000 output tokens remains in the lower tier; a request with 120,000 input tokens uses the higher input and output rates listed for the long-prompt tier.

For a simple request with 10,000 input tokens and 500 output tokens, the lower-tier arithmetic is:

(10,000 x $0.10 + 500 x $0.50) / 1,000,000 = $0.00035

At 10,000 identical requests, that is $3.50 before tool, platform, or retry costs. To show the tier change with a valid long prompt, a 120,000-input-token request with 500 output tokens costs:

(120,000 x $0.50 + 500 x $2.50) / 1,000,000 = $0.06125

That is $612.50 for 10,000 such requests. Count the complete input sent to the API, not only the visible question.

Anthropic's official Reddit announcement says requests up to 100K tokens are about 90% cheaper than Haiku 4.5 and requests above 100K are about 50% cheaper. Treat those as relative launch claims, not as a guarantee that every workload will cost those percentages less: retries, output length, tokenizer behavior, and cache misses still change the bill.

Caching and Batch change the effective rate

Prompt caching is particularly useful when an application sends the same instructions or knowledge pack repeatedly. Haiku 5.5's five-minute cache write is 1.25 times the standard input rate, while a cache read is one-tenth of standard input pricing. The higher prompt tier keeps the same ratios.

OperationUp to 100KOver 100K
First 1M tokens written to five-minute cache$0.125$0.625
Each 1M cached tokens read$0.01$0.05
Each 1M input tokens in Batch$0.05$0.25
Each 1M output tokens in Batch$0.25$1.25

A five-minute cache usually breaks even after roughly two reads because the write premium is recovered quickly. For example, a 1M-token prefix costs $0.125 to write and $0.01 to read, versus $0.20 for two uncached reads, before output. A one-hour cache is priced at twice the base input rate, so it needs more repeated reads to pay back. The exact result depends on whether the cache key survives unchanged and whether the request remains in the same pricing tier.

Batch API halves input and output rates for asynchronous jobs such as nightly enrichment, bulk classification, backfills, and evaluation runs. It is not a substitute for synchronous chat: latency is the trade-off. If your route supports both, Batch and caching can be combined; validate the final usage fields rather than assuming every request hit the cache.

The operational signal to log is cache_read_input_tokens, a field documented in Anthropic's Messages API usage schema. If it stays at zero after the first request, inspect prompt ordering, timestamps, tool serialization, and cache-control placement. A cheaper model with a perpetually cold cache can cost more than a correctly cached route with a higher nominal rate.

Haiku 5.5 versus older Haiku and Sonnet

The comparison below uses the published direct rates available before Haiku 5.5. Haiku 4.5's rates are included as a baseline; they are not a claim that every provider has identical pricing.

ModelInput / MTokOutput / MTokBest budget interpretation
Claude Haiku 5.5, <=100K$0.10$0.50High-volume, short-output work
Claude Haiku 5.5, >100K$0.50$2.50Long prompts where quality justifies the tier
Claude Haiku 4.5$1.00$5.00Existing integrations needing a known baseline
Claude Sonnet 5.5$2.00$10.00More demanding production tasks

Haiku 5.5 is the obvious first test for title generation, classification, extraction, routing, short summaries, and other high-throughput worker calls. A real user on Reddit described Haiku-style use as “high-volume, low-complexity tasks” such as creating a conversation title or extracting software names from text (u/jam_pod_). That is a better fit than using it as an assumed replacement for a stronger model on every agent step.

Do not choose by token price alone. If Haiku needs a second attempt on half of your requests while Sonnet produces accepted results on the first attempt, the nominal two-to-one rate gap can disappear. Measure billed input, billed output, retries, and accepted results on a representative sample before routing production traffic.

A production cost rule

Use this sequence before committing to Haiku 5.5:

  1. Run the real prompt through Anthropic's token-counting documentation, including tools, retrieved documents, and conversation history.
  2. Split traffic into <=100K and >100K input requests; do not average the tiers together.
  3. Estimate output separately. A short input with a long generated answer is output-dominated.
  4. Put stable prefixes behind cache control and alert when cache reads fall to zero.
  5. Send work that can wait through Batch API and retain synchronous routing for interactive requests.
  6. Compare Haiku 5.5 with Sonnet 5.5 by cost per accepted result, not cost per million tokens.

This rule also handles the main uncertainty in the launch-day data: the official pricing and availability are clear, but independent latency, failure-rate, and long-context quality measurements are not yet established. The Reddit discussion contains the same unresolved question in plain terms: “What does ‘prompts up to 100k’ mean?” (u/ShadowBannedAugustus). For a production team, that question belongs in a billing test, not a guess.

Claude Haiku 5.5 API FAQ

What is the Claude Haiku 5.5 API model ID?

Use claude-haiku-5-5 as the current alias, as listed in Anthropic's Claude Code release notes. Confirm the snapshot ID in the API model documentation when reproducible routing matters.

How much does a 100K-token Haiku 5.5 request cost?

For 100,000 input tokens and no output, the lower-tier input charge is $0.01. Output is billed separately at $0.50 per million tokens in the same tier; the higher tier applies when the prompt exceeds the documented boundary.

Is Haiku 5.5 included in Claude Pro or Max?

A Claude subscription and Claude API usage are separate billing products, as shown in Anthropic's Claude platform pricing. Use the API rate card for programmatic traffic.

Does Batch pricing stack with prompt caching?

Anthropic's pricing documentation presents Batch as a 50% discount on eligible token charges, and cached-token accounting remains separate. Validate the final usage fields for your route because cache eligibility and provider-specific support can differ.

Is Haiku 5.5 cheaper than Haiku 4.5 for every request?

At the published direct rates, Haiku 5.5 is cheaper in both input tiers. The practical saving can be reduced by retries, longer outputs, cache misses, or a workload that needs a stronger model to finish successfully.