AIREITER

Claude API Pricing 2026: Every Model & How to Pay Less

Last Updated: 2026-06-30 08:51:38

Claude API pricing runs from $1 per million tokens (Haiku 4.5 input) to $50 per million (Fable 5 output) as of July 2026, and across every current model the rule holds: output tokens cost 5× input tokens. But the sticker price is the easy part. A new tokenizer, US-only routing, fast mode, and a billing trap most people never notice all sit between the rate card and your actual invoice — and there's a way to buy the same models for a fraction of list price.

Every number below is on Anthropic's official pricing page; check it before you finalize a budget, since model rates move.

Claude API pricing in 2026: every model

These are the current per-million-token rates, verified against Anthropic's pricing for July 2026 (official source). Opus 4.8 launched May 28, 2026 at the same rate as 4.7.

Model

Input / 1M

Output / 1M

Context

Claude Fable 5

$10.00

$50.00

1M

Claude Opus 4.8

$5.00

$25.00

1M

Claude Opus 4.7

$5.00

$25.00

1M

Claude Opus 4.6

$5.00

$25.00

1M

Claude Sonnet 4.6

$3.00

$15.00

1M

Claude Haiku 4.5

$1.00

$5.00

200K

Two things a lot of older guides get wrong. Fable 5 — Anthropic's most capable model — sits above the Opus tier at $10/$50, and the whole Opus 4.x line (4.6, 4.7, 4.8) holds at $5/$25. That $5 is itself a story: Opus 4.1 cost $15/$75, and the drop to $5/$25 at the 4.6 launch was a 67% cut on the flagship that has held through three releases since.

Legacy and budget options

If you don't need the newest models, the older ones are still live: Sonnet 4.5 ($3/$15), Opus 4.5 ($5/$25, 200K context), Haiku 3.5 ($0.80/$4), and Haiku 3 ($0.25/$1.25) for the most cost-sensitive, high-volume work.

One underrated detail: Opus 4.6/4.7/4.8, Sonnet 4.6, and Fable 5 all include the full 1M-token context window at standard pricing — no long-context surcharge. A 900K-token request costs the same per-token rate as a 9K one, where several competitors charge a premium past a length threshold.

The costs that aren't on the price card

The rate table is where most pricing articles stop, and where the surprises start — because several things move your real bill without changing the per-token number.

The tokenizer that quietly adds tokens

Opus 4.7 (and Opus 4.8, which shares its tokenizer) changed how text is split into tokens. The same input can produce up to ~35% more tokens than on Opus 4.6 — at the identical per-token rate. Code, structured data (JSON/XML), and non-English text are hit hardest, because the rate didn't change but the token count did. Don't take the 35% on faith: before you trust an old budget, re-run your actual prompts through the token-counting endpoint on the new model and compare.

Thinking, fast mode, and US-only routing

Per Anthropic's pricing docs:

  • Extended thinking tokens are billed as output, even when you never display them. A reasoning-heavy request quietly runs up the output side of the bill.

  • Fast mode (Opus 4.8/4.7 only) runs the same model up to 2.5× faster at 2× standard pricing.

  • US-only inference (data residency) adds a 1.1× multiplier on input and output; AWS Bedrock or Vertex regional endpoints add roughly another 10% on top.

Server-side tools

Tools that run on Anthropic's side bill separately, per the same pricing docs: web search is $10 per 1,000 searches plus the tokens to process results, and code execution is free up to a monthly allowance, then about $0.05/hour per container. (Confirm the current figures on the pricing page — server-tool rates change.)

The billing trap that catches Claude Code users

This one isn't on any price card. If you have ANTHROPIC_API_KEY set in your shell, Claude Code bills through the pay-per-token API instead of your subscription. People expecting a flat $20–$200/month have been blindsided by metered charges for exactly this reason. Check your environment before a long session.

How much will Claude actually cost you?

Rates mean little until you map them to a workload. Three examples at current prices, all at 10,000 requests/day (2,000 input + 500 output tokens each — 20M input + 5M output tokens/day):

  • A support chatbot on Haiku 4.5 ($1/$5): 20M × $1 + 5M × $5 = $45/day, about $1,350/month.

  • A coding agent on Sonnet 4.6 ($3/$15): $60 + $75 = $135/day, about $4,050/month before optimization.

  • A high-context RAG app on Opus 4.8 ($5/$25), where retrieved documents push input far higher, climbs into the high four or five figures a month because input volume dominates.

Same volume, three models — the model choice alone swings the bill ~3× between Haiku and Sonnet.

These aren't hypothetical ceilings. One r/ClaudeAI thread tracked 31 Claude Code users whose usage would have cost roughly $80,000/month at API rates, the top user alone near $18,000 — an estimate of what their subscription usage would bill on the metered API, not a literal invoice. Heavy autonomous workloads are where the bill gets real, which is exactly where the optimizations below pay off.

Rate limits and usage tiers

Pricing isn't the only constraint — how fast you can spend it is tiered. New API accounts start with conservative requests-per-minute (RPM) and tokens-per-minute (TPM) limits that rise automatically as your account ages and your cumulative spend grows (Tier 1 → 4). New accounts also get $5 in free credits to test with, no card required. Beyond throughput, Anthropic offers service tiers that trade off availability and price: a standard tier (the default), a priority tier for guaranteed availability, and a batch tier for the 50% async discount below. Check the rate-limits docs for your tier's exact numbers before you architect for scale — a quota that served pilot traffic may need a tier bump for production.

Claude subscription vs API: which is cheaper?

The most common real-world question, and the answer is a threshold, not a verdict.

  • Claude Pro is $20/month; Max is $100 or $200/month. These cover the Claude apps and Claude Code under a usage allowance — no per-token billing.

  • The API is pure pay-as-you-go, no monthly minimum, no usage ceiling.

The crossover sits in the low millions of tokens per month — roughly 3–4M on a mid-tier model (Sonnet), assuming a balanced input/output mix and that you'd otherwise stay inside Pro's usage allowance. It shifts up on Haiku and down on Opus, so treat it as a range, not a fixed line. Below that volume a subscription is usually cheaper and simpler. Above it — especially production traffic or heavy agentic coding — the API's per-token rate wins, and you gain control (rate limits, monitoring, multi-key billing) a plan doesn't give. For all-day coding specifically, developers estimate a Max plan can be dramatically cheaper than the same work billed token-by-token on Opus.

The practical rule: prototypes and individual use → subscription; scaled product traffic → API. And mind the ANTHROPIC_API_KEY trap above, or you'll pay both ways.

How to pay less for the Claude API

Three levers cut the standard bill without changing models.

1. Prompt caching (up to 90% off repeated context)

Cache a stable prefix — a long system prompt, retrieved documents, few-shot examples — and cache reads cost 0.1× the input rate, a 90% discount. The catch is the write: a 5-minute cache write costs 1.25× and pays for itself after one read; a 1-hour write costs 2× and pays off after two. For anything with a large fixed preamble hit repeatedly, this is the biggest single lever. (Caching docs.)

2. Batch API (50% off, both directions)

For anything that doesn't need a real-time answer, the Batch API discounts input and output by 50% and returns within 24 hours. Sonnet 4.6 drops to $1.50/$7.50; Haiku 4.5 to $0.50/$2.50. Caching and batching stack, pushing effective cost well below half the sticker rate.

3. Model routing and output caps

Don't run every request on Opus. Route classification, summarization, and routing logic to Haiku, reserve Sonnet for the middle, and call Opus only when the task needs it — commonly a 40–60% cut on its own. And because output is 5× the price of input, capping max_tokens is one of your highest-leverage knobs: trimming average output from 800 to 500 tokens across a million calls saves real money every month.

Official API vs relay platforms: can you pay less?

There's a fourth route the official docs won't mention: buying the same Claude models through a relay. An API relay (or aggregator) sits between you and Anthropic, exposes an Anthropic-compatible endpoint, and charges its own rate.

Why can a relay be cheaper? Three mechanics, mostly: relays buy at committed-use volume that earns discounts an individual account can't reach; they route efficiently across regions and providers; and many run thin margins (or a loss-leader) to win developers. But a discount that deep has to come from somewhere — understand a platform's model before you depend on it.

Discounts vary widely. Here's how three options compare for the same Claude models:

Platform

Discount vs Anthropic list

Model coverage

Notes

Official Anthropic API

— (baseline)

All, incl. Fable 5 / Opus 4.8

Caching, batch, every feature

OpenRouter

≈ list (passes provider price through)

Many models, multi-provider

Aggregator; convenience over savings

EvoLink

~10% off

Up to Opus 4.7

OpenAI-compatible endpoint

AIReiter

~80% off

Up to Opus 4.6

Anthropic-compatible; caching passes through

AIReiter sits at the aggressive end. Here's what the headline ~80% works out to per model:

Model

Official (in / out)

AIReiter (in / out)

You save

Claude Haiku 4.5

$1 / $5

$0.20 / $1.00

80%

Claude Sonnet 4.6

$3 / $15

$0.60 / $3.00

80%

Claude Opus 4.6

$5 / $25

$1.00 / $5.00

80%

Integration is drop-in: AIReiter's endpoint is compatible with the Anthropic Messages API, so the official anthropic SDK works by changing two lines — point base_url to https://aireiter.com/api and swap the key:

client = anthropic.Anthropic(
    api_key="YOUR_AIREITER_KEY",
    base_url="https://aireiter.com/api",
)

Streaming, multi-turn, vision, and cache_control all pass through, so the 90% caching discount above stacks on top of the lower base rate.

FAQ

Why is Claude API so expensive?

Often it isn't the per-token rate — it's workflow design. The biggest cost drivers are running everything on Opus when Haiku or Sonnet would do, uncached repeated context, uncapped output, and reasoning tokens you're billed for but never see. Fixing those usually matters more than the headline price.

Is there a free Claude API tier?

No permanent free production tier, but new accounts get $5 in free credits (no card required) to test with. After that it's pay-as-you-go.

Claude API vs ChatGPT pricing — which is cheaper?

It depends on the tier you compare. Roughly, per million tokens (in / out):

Claude

Comparable OpenAI

Haiku 4.5 — $1 / $5

GPT mini tier — low single digits

Sonnet 4.6 — $3 / $15

mid tier — comparable

Opus 4.8 — $5 / $25

GPT-5.5 — ~$5 / $30

How do I calculate my monthly Claude API bill?

Estimate it as: (requests × avg input tokens ÷ 1,000,000 × input rate) + (requests × avg output tokens ÷ 1,000,000 × output rate). Use the token-counting endpoint for accurate counts on your real prompts — a generic tokenizer can be off by 15–35% on Claude.

The bottom line

For most teams it's a three-way split: prototyping or individual use → a Pro or Max subscription (simpler, cheaper below ~3–4M tokens/month); production on the latest flagship (Opus 4.8, Fable 5) → the official API with caching, batching, and routing on; production where Sonnet 4.6 or Opus 4.6 is enough → a reputable relay like AIReiter can cut the same bill by up to 80% with a drop-in SDK swap.

Whichever you pick, most teams overpay by skipping the optimizations, not by choosing the wrong tier — so turn the levers on before you worry about the rate card.