A Claude Code user priced one month of Max-plan work at API list rates and landed on $7,470 against a $200 invoice. The per-token rate card is not the anomaly. Metered access removes a ceiling the subscription was quietly enforcing, and agent workloads resend far more tokens than a human typing into a chat box ever will.
The gap developers report
The ratios people post are large and fairly consistent. A r/ClaudeAI thread reconstructing Claude Code session logs put one month of Max usage at roughly $7,470 of API list price against a $200 subscription. In a smaller subscription-vs-API thread, a $200 plan holder priced the same workload at $5,000–$20,000 per month.
Those are unaudited user estimates, and in both the token-heaviest pattern is the same: an agent looping over a repository.
"I'm on enterprise API plan for my company via claude code and I can tell you first hand that you can easily burn $100-$200 per session in API costs just coding normally on larger projects." — u/andywidjaja, r/ClaudeAI
A paid plan and an API key are different products
Anthropic states the separation directly. Its help-center article, updated March 16, 2026, calls Claude plans and the Claude Console "separate products designed for different purposes" and says a paid subscription "doesn't include access to the Claude API or Console." OpenAI runs the same split, billing ChatGPT and the API platform through separate systems.
What each side sells is the better lens:
| Paid chat plan | API key | |
|---|---|---|
| Unit sold | One human seat | Token throughput |
| Ceiling | Rolling 5-hour window plus weekly limits | No bundled quota; rate and spend limits still apply |
| Delivery | Web, desktop, mobile apps | Your own code |
| Price shape | Fixed monthly | Variable, per token |
| Powering your own product | Not what it is sold for | That is the point |
Claude's pricing page lists Pro at $20/month monthly or $17/month on annual billing ($200 up front), and Max from $100/month at either 5× or 20× Pro usage per five-hour session. Anthropic publishes no message count for those tiers, because consumption varies with conversation length, model, and features, so the comparison pits a measured quantity against an unmeasured one.
The same page carries one more tell: Claude Enterprise pairs a $20 per-seat monthly fee with usage billed at API rates, so heavy customers land back on the meter.
What a subscription dollar buys at list price
Converting plan dollars into tokens makes the gap concrete. These are official rates per million tokens (MTok) from Claude's platform pricing docs and OpenAI's API pricing page, checked October 1, 2026.
| Model | Input | Output | Cache read | Batch (in/out) |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $0.25 | $5 / $25 |
| Claude Opus 5.5 | $4 | $20 | $0.20 | $2 / $10 |
| Claude Sonnet 5.5 | $2 | $10 | $0.20 | $1 / $5 |
| Claude Haiku 4.5 | $1 | $5 | $0.10 | $0.50 / $2.50 |
| GPT-6 Astra | $10 | $50 | $1.00 | $5 / $25 |
| GPT-6.1 Sol | $2 | $10 | $0.10 | $1 / $5 |
| GPT-6 Luna | $0.10 | $0.50 | $0.01 | $0.05 / $0.25 |
For a chat-shaped workload, a few thousand words of context and a few hundred words back, $20 of Sonnet 5.5 credit covers roughly 10 million input tokens or 2 million output tokens. Your own number comes from the four fields in the API response: (fresh input × input rate) + (cached input × cache-read rate) + (cache writes × write rate) + (output × output rate), with batch, long-context, and region multipliers on top. Log those four per endpoint for a week and you have a monthly figure to compare against a seat.
Where the tokens actually go
Four common sources of unexpected token spend sit off the rate card. The last three are documented in LLM Cost Lab's breakdown of off-rate-card costs.
- The transcript, resent every turn. In request-based chat integrations the client resends history, so turn 20 carries turns 1 through 19 with it. When each turn adds a similar amount of context, input volume grows roughly with the square of conversation length, which is why that breakdown prices a 20-turn session at $0.163 against $0.063 budgeted as independent calls, a 2.6× gap.
- The system prompt, multiplied by every call. A 2,000-token system prompt across 1 million requests is 2 billion input tokens before a user types anything.
- Reasoning tokens you never see. On models that bill hidden thinking traces as output, those tokens can run several times longer than the visible answer, so metering on returned characters undercounts the invoice.
- Tool results, retries, and discarded generations. If the model generated it, you pay, including the response your JSON validator rejected and the parallel call you raced and threw away.
Cheap tokens can still dominate a bill, which is what confuses people reading their own usage dashboard. On the $7,470 thread, a commenter noticed cache reads made up 48% of the total despite being the cheapest token type on the card.
"that's already the cheapest token type on the api, so if it's still dominant after applying list prices, the actual workload is mostly re-reading context every turn" — u/YoanEdwin, r/ClaudeAI
Put Opus 5.5 rates on an agent turn carrying 150k tokens of repo context and the structure becomes visible:
| Opus 5.5, 60-turn session, 150k context + 2k output per turn | Cost |
|---|---|
| Fresh input per turn, no cache hit (150k @ $4) | $0.60 |
| Cache read per turn instead of fresh input (150k @ $0.20) | $0.03 |
| Output per turn (2k @ $20) | $0.04 |
| One 1-hour cache write per session (150k @ $8) | $1.20 |
| Session total, uncached | $38.40 |
| Session total, cached | $5.40 |
Two cached sessions a day across 22 working days is roughly $238, already past a $200 plan. Uncached, the same month clears $1,600. The rate never changed; the token count did.
Three rungs, not two
Most comparisons stop at plan versus API. A third rung exists, and users meet it only after a quota runs dry: extra usage credits sold inside the subscription product.
Reports from r/codex put that rung above raw API pricing. One Pro user tracked roughly 16M tokens consumed through purchased credits costing about $40, then estimated the same volume at roughly $12 on list API rates.
"yeah the credit pricing is borderline predatory. $50 for a single landing page is absurd when the api would cost like $2." — u/AbjectBug5885, r/codex
These are user estimates, not universal rates, and top-up pricing differs by vendor and plan. In the Codex cases above, the top-up rung priced roughly 3× the equivalent list API spend, which is worth checking against your own numbers before buying credits a second week running.
Five levers that move the number
Each lever carries an official rate, so the arithmetic is checkable against the vendor page first.
- Prompt caching, with the write premium included. Anthropic's pricing docs put a 5-minute cache write at 1.25× base input and a 1-hour write at 2×, with reads at 0.1× base (0.05× on Opus 5.5, 0.025× on Fable 5.1). A 5-minute write pays for itself after one read, a 1-hour write after two. Written once and never read, it costs more than not caching at all. Ordering matters as much as enabling it: stable content first, user variation last, or the prefix stops matching. Hit-rate debugging is covered in this prompt caching guide.
- Batch anything that can wait. Both vendors cut list price 50% for asynchronous processing: Opus 5.5 drops to $2/$10, Sonnet 5.5 to $1/$5. Coverage is uneven elsewhere: LLM Cost Lab counts 26 of 83 tracked models exposing a discounted batch tier, so confirm yours does before architecting around it. Mechanics are in our batch API pricing walkthrough.
- Tier the model to the task. Classification, routing, and retrieval decisions rarely need a frontier model. Haiku 4.5 at $1/$5 against Fable 5.1 at $10/$50 is a 10× spread on steps that emit ten tokens.
- Constrain the output, not just
max_tokens. Output is priced 5× input across nearly every row in the table above. A schema that forbids preamble trades a little structural overhead for the prose you would otherwise be billed for, andmax_tokensworks better as a safety ceiling than as a formatting tool, since a truncated answer may require a second paid attempt. - Watch the multipliers riding on top. OpenAI's long-context tier doubles short-context input rates, and Fast mode doubles standard pricing again. Anthropic charges 1.1× for US-pinned inference, and its pricing docs state that Claude 4.7 and later use a newer tokenizer producing roughly 30% more tokens for the same text. Nothing about your prompt changed; the meter did.
One structural option sits outside that list: route programmatic traffic through a single metered endpoint rather than one billing relationship per vendor, which is what AIReiter's Claude API endpoint is for. At list rates, a month of overflow at 40M Sonnet-class input tokens and 8M output is $80 + $80, so the comparison against a second seat takes about a minute.
Which one to buy
| Your situation | Buy | Why |
|---|---|---|
| Chat, research, occasional code in an app | Paid plan only | Capped spend, bundled apps, usage well under plan limits |
| Solo agentic coding, one machine | Plan first, metered key for overflow | The plan absorbs the bulk; the key removes the weekly-reset wait |
| Shipping a product other people use | API only | A seat covers one human; your users need their own throughput |
| Periodic bulk jobs: evals, backfills, classification | API with batch tier | 50% off list, latency irrelevant |
| Already buying top-up credits most weeks | Price the overflow against a key | Reported credit pricing sat above list API rates |
FAQ
Does a paid ChatGPT or Claude plan include API credits?
No. Anthropic's help center states plainly that a paid Claude plan "doesn't include access to the Claude API or Console," and OpenAI bills ChatGPT and the API platform through separate systems. Paying for both means two line items.
Am I billed for reasoning tokens I never see?
Yes, on models that expose thinking or extended reasoning. Those tokens bill at the output rate without appearing in the response body, so read the usage object.
Do I pay for the whole conversation on every turn?
Yes, unless a cache hit covers the repeated prefix. Absent a server-side state feature, your client resends the history it wants the model to see, and that history is billed as fresh input at the model's rate, or at the cache-read rate when the prefix matched.
Does a failed or retried request still cost money?
Per LLM Cost Lab's billing breakdown, a request that reaches the model and returns is billed even when your application discards the result after a schema failure or client timeout. Requests rejected before generation are the exception.
The trade-off nobody has resolved
Plans cap spend but ration access; metered keys invert both. With no published token value for a plan quota, answering "which is cheaper for me" takes a week of your own instrumented traffic priced against the table above.
Related reading: Cheapest LLM API in 2026 · Claude API pricing guide