AIREITER

DeepSeek API Price Increase: Old vs New Rates (Aug 2026)

Last Updated: 2026-08-16 18:58:05

DeepSeek's API price increase took effect August 16, 2026, replacing flat per-token rates with peak/off-peak billing on both deepseek-v4-flash and deepseek-v4-pro. Cache-miss input and output now cost 1.5× to 2.4× the old prices off-peak, double that at peak; cache-hit input rose 2.5× to 6.1× off-peak and up to 12.1× at peak. The number that decides how bad this is for you: your prompt-cache hit rate.

What changed on August 16, 2026

The switch hit at 16:00 UTC on Sunday August 16 - midnight August 17 Beijing time - per both the announcement thread and the 650-call audit on r/DeepSeek. The official Models & Pricing page now lists the new tables, so the change is live, not announced-only.

DeepSeek split its billing into peak and off-peak windows, with peak rates set at exactly 2× off-peak rates. The two peak windows are 01:00–04:00 UTC and 06:00–10:00 UTC - seven peak hours per day, seventeen off-peak. The model versions (DeepSeek-V4-Flash-0731, DeepSeek-V4-Pro-0813), API compatibility, and published specs are unchanged on that same page - this is a billing change, not a model change.

The new DeepSeek API pricing tables

All prices below are USD per 1M tokens, straight from the official pricing page as of August 16, 2026.

V4 FlashOff-peakPeak
Input, cache hit$0.007$0.014
Input, cache miss$0.22$0.44
Output$0.66$1.32
V4 ProOff-peakPeak
Input, cache hit$0.022$0.044
Input, cache miss$0.66$1.32
Output$1.98$3.96
DeepSeek official API pricing page showing new peak and off-peak rates, August 16, 2026

Billing mechanics are otherwise unchanged: requests are billed at the rate in effect when they run, and DeepSeek explicitly reserves the right to change prices again.

How much did the DeepSeek API price increase actually raise rates?

The old flat rates were verified by third-party trackers through July 31, 2026: Flash at $0.14 input (cache miss), $0.0028 cache hit, $0.28 output; Pro at $0.435 / $0.003625 / $0.87 (TLDL, BenchLM). Running both tables side by side:

Per 1M tokensOld flatNew off-peakNew peakOff-peak vs oldPeak vs old
Flash input, cache hit$0.0028$0.007$0.0142.5×5.0×
Flash input, cache miss$0.14$0.22$0.441.57×3.14×
Flash output$0.28$0.66$1.322.36×4.71×
Pro input, cache hit$0.003625$0.022$0.0446.07×12.14×
Pro input, cache miss$0.435$0.66$1.321.52×3.03×
Pro output$0.87$1.98$3.962.28×4.55×
DeepSeek V4 API prices: old flat rate vs new off-peak vs new peak, USD per million tokens

The headline "1,114% increase" circulating on Reddit is real but narrow: it's V4 Pro cache-hit input at peak ($0.003625 -> $0.044). A zero-cache workload off-peak sees closer to 1.5–2.3×. The multiplier you land on between those two extremes is the whole story.

Why cache-heavy workloads get hit hardest

Cache-hit prices rose 2.5× to 12× while cache-miss and output prices rose 1.5× to 4.7× - so the more of your input volume that was riding the ultra-cheap cache-hit lane, the bigger your percentage jump. One r/DeepSeek user repriced 650 real API calls (16 sessions, 155.9M prompt tokens, 98.09% cache-hit rate) through both price tables and reconciled the result against DeepSeek's console billing within 2%:

"the better caching works for you today, the bigger your increase" - u/Swimming-Soup-1173, r/DeepSeek audit of 650 API calls

Their agent workload repriced at 2.7× off-peak, versus 1.5× for the same workload with zero cache hits. Another user modeled a long-context workload going from $33 to $100 off-peak or $200 at peak - a 3× to 6× jump. Batch jobs, RAG pipelines, and agents with static system prompts are the profiles getting squeezed; one-shot summarization traffic sits closer to the 1.5–2.3× off-peak floor by comparison.

Peak hours translated to your timezone

The official page only lists UTC windows. Converted for August (daylight saving in effect):

ZonePeak windows (local)Your 9-to-5
Beijing (UTC+8)09:00–12:00, 14:00–18:00Peak except the 12:00–14:00 lunch
Central Europe (UTC+2)03:00–06:00, 08:00–12:00Mornings peak
London (UTC+1)02:00–05:00, 07:00–11:00Mornings peak
US East (UTC-4)21:00–00:00, 02:00–06:00Fully off-peak
US West (UTC-7)18:00–21:00, 23:00–03:00Fully off-peak

The peak windows cover the Chinese business day minus the lunch hour (12:00–14:00 Beijing time is off-peak). US teams run their whole workday at off-peak rates, while European teams pay double on morning batches (08:00–12:00 CEST is peak) and should shift that work to the afternoon or overnight.

Is prompt caching still worth it?

Yes, unambiguously - the discount shrank, it didn't disappear. Cache-hit input was previously ~50× cheaper than cache-miss for Flash (120× for Pro); it is now ~30× cheaper for both models (31× for Flash, 30× for Pro off-peak: $0.007 vs $0.22, $0.022 vs $0.66). The audit author who measured the 2.7× jump still concluded:

"Caching is still absolutely worth it" - same r/DeepSeek audit

In their workload, caching still cut costs roughly 15× versus sending the same context uncached. Caching remains best-effort, with no guaranteed hit rate, and unused entries are generally cleared after hours to days (BenchLM) - so measure before budgeting, logging the usage object's prompt_cache_hit_tokens and prompt_cache_miss_tokens fields as BenchLM recommends to compute your actual hit share.

Is DeepSeek still the cheap option after the increase?

Off-peak, yes - at peak, the margin over GPT-5.6 narrows to nothing on the Flash tier. Output price per 1M tokens, standard rates:

ModelInputOutputNotes
DeepSeek V4 Flash (off-peak)$0.22$0.661M context
DeepSeek V4 Flash (peak)$0.44$1.32Exceeds Luna output price
GPT-5.6 Luna$0.20$1.20Price cut 80% on July 30
DeepSeek V4 Pro (off-peak)$0.66$1.98Still ~6× under Terra
DeepSeek V4 Pro (peak)$1.32$3.96~3× under Terra
GPT-5.6 Terra$2$12
Claude Sonnet 5$2$10Intro rate through Aug 31, then $3/$15

Off-peak, V4 Flash output at $0.66 undercuts GPT-5.6 Luna by 45% and Terra by 18× - DeepSeek stays cheapest on output for high-volume work. At peak, Luna ($1.20 output, $0.20 input) beats Flash on both standard rates, though Luna's cache-read rate ($0.02/M in the AI Price Index changelog) is pricier than Flash's cache-hit lane at any hour, and Flash still lists a 1M-token context and 2,500 concurrency on the official specs table. As one commenter in the announcement thread put it: "DS4 is good but when cheap."

What to do this week

  1. Reprice your own bill before panicking. Pull last month's usage logs and compute: cache-hit input × hit rate × hit price + cache-miss input × miss rate × miss price + output × output price, applying the peak or off-peak rate in force for each request's window. The audit's endpoints bracket the range: zero cache hits repriced at 1.5×, a 98% hit rate at 2.7×.
  2. Shift batchable work off peak. Cron jobs, evaluations, document processing - anything without a user waiting can run in the 17 off-peak hours. For US teams this is nearly free (your workday already is off-peak); for EU teams, move morning batches to the afternoon.
  3. Keep producing cache hits. Whatever prompt structure is already generating cache hits for you, keep it - the cache-hit lane still runs ~30× below cache-miss input.
  4. Test a Flash downgrade for routine calls. Pro runs 3× Flash across the board (3.1× on cache hits). If you haven't re-tested Flash since the 0731 post-training update, route a slice of routine traffic to it - the model-selection tradeoffs are covered in our V4 Flash vs V4 Pro comparison.
  5. If you leave, migration can be small. For OpenAI-compatible clients, leaving may only mean changing the base URL and model ID. Third-party hosts and relay platforms bill on their own rate cards - DataLLM Lab documented hosts quoting V4 Pro around $1.74/$3.48 per 1M even before this increase - so compare the effective rate, not the sticker.

The decision in one table:

Your profileCall
US timezone, cache-heavy agent trafficStay - off-peak rates plus the ~30× cache lane keep input cost lowest for cache-heavy traffic
EU timezone, morning batch jobsStay, reschedule - peak exposure is self-inflicted and avoidable
Peak-hours-only, quality-sensitive coding on ProRe-evaluate - Pro peak output ($3.96) is still under Terra ($12), but the gap is now 3×, not the 14× it was before August 16
Zero-cache, one-shot workloadsStay - 1.5–2.3× off-peak is the smallest increase in the table

FAQ

When exactly did the new DeepSeek prices take effect?

August 16, 2026 at 16:00 UTC (midnight August 17 Beijing time), per both the announcement thread and the 650-call audit on r/DeepSeek. The official pricing page carried the new tables that same day.

How much will my DeepSeek bill increase?

A zero-cache workload reprices at roughly 1.5× off-peak and a measured 98%-hit-rate workload at 2.7×; peak hours double those figures, up to 12× on Pro cache-hit input. The multiplier table above has every rate pair.

Which hours are peak in US time?

US Eastern: 21:00–00:00 and 02:00–06:00. US Pacific: 18:00–21:00 and 23:00–03:00. Normal US business hours fall entirely in off-peak windows.

Are deepseek-chat and deepseek-reasoner still available?

No. The legacy aliases reached their retirement deadline on July 24, 2026, per BenchLM's model-ID history; new integrations should call deepseek-v4-flash or deepseek-v4-pro.

Is DeepSeek still cheaper than GPT-5.6 and Claude?

Off-peak, yes - Flash output is 45% under GPT-5.6 Luna and 18× under Terra. At peak, Luna's $1.20/M output undercuts Flash's $1.32/M, so the answer depends on when your traffic runs.