DeepSeek's API price increase took effect August 16, 2026, replacing flat per-token rates with peak/off-peak billing on both deepseek-v4-flash and deepseek-v4-pro. Cache-miss input and output now cost 1.5× to 2.4× the old prices off-peak, double that at peak; cache-hit input rose 2.5× to 6.1× off-peak and up to 12.1× at peak. The number that decides how bad this is for you: your prompt-cache hit rate.
What changed on August 16, 2026
The switch hit at 16:00 UTC on Sunday August 16 - midnight August 17 Beijing time - per both the announcement thread and the 650-call audit on r/DeepSeek. The official Models & Pricing page now lists the new tables, so the change is live, not announced-only.
DeepSeek split its billing into peak and off-peak windows, with peak rates set at exactly 2× off-peak rates. The two peak windows are 01:00–04:00 UTC and 06:00–10:00 UTC - seven peak hours per day, seventeen off-peak. The model versions (DeepSeek-V4-Flash-0731, DeepSeek-V4-Pro-0813), API compatibility, and published specs are unchanged on that same page - this is a billing change, not a model change.
The new DeepSeek API pricing tables
All prices below are USD per 1M tokens, straight from the official pricing page as of August 16, 2026.
| V4 Flash | Off-peak | Peak |
|---|---|---|
| Input, cache hit | $0.007 | $0.014 |
| Input, cache miss | $0.22 | $0.44 |
| Output | $0.66 | $1.32 |
| V4 Pro | Off-peak | Peak |
|---|---|---|
| Input, cache hit | $0.022 | $0.044 |
| Input, cache miss | $0.66 | $1.32 |
| Output | $1.98 | $3.96 |
Billing mechanics are otherwise unchanged: requests are billed at the rate in effect when they run, and DeepSeek explicitly reserves the right to change prices again.
How much did the DeepSeek API price increase actually raise rates?
The old flat rates were verified by third-party trackers through July 31, 2026: Flash at $0.14 input (cache miss), $0.0028 cache hit, $0.28 output; Pro at $0.435 / $0.003625 / $0.87 (TLDL, BenchLM). Running both tables side by side:
| Per 1M tokens | Old flat | New off-peak | New peak | Off-peak vs old | Peak vs old |
|---|---|---|---|---|---|
| Flash input, cache hit | $0.0028 | $0.007 | $0.014 | 2.5× | 5.0× |
| Flash input, cache miss | $0.14 | $0.22 | $0.44 | 1.57× | 3.14× |
| Flash output | $0.28 | $0.66 | $1.32 | 2.36× | 4.71× |
| Pro input, cache hit | $0.003625 | $0.022 | $0.044 | 6.07× | 12.14× |
| Pro input, cache miss | $0.435 | $0.66 | $1.32 | 1.52× | 3.03× |
| Pro output | $0.87 | $1.98 | $3.96 | 2.28× | 4.55× |
The headline "1,114% increase" circulating on Reddit is real but narrow: it's V4 Pro cache-hit input at peak ($0.003625 -> $0.044). A zero-cache workload off-peak sees closer to 1.5–2.3×. The multiplier you land on between those two extremes is the whole story.
Why cache-heavy workloads get hit hardest
Cache-hit prices rose 2.5× to 12× while cache-miss and output prices rose 1.5× to 4.7× - so the more of your input volume that was riding the ultra-cheap cache-hit lane, the bigger your percentage jump. One r/DeepSeek user repriced 650 real API calls (16 sessions, 155.9M prompt tokens, 98.09% cache-hit rate) through both price tables and reconciled the result against DeepSeek's console billing within 2%:
"the better caching works for you today, the bigger your increase" - u/Swimming-Soup-1173, r/DeepSeek audit of 650 API calls
Their agent workload repriced at 2.7× off-peak, versus 1.5× for the same workload with zero cache hits. Another user modeled a long-context workload going from $33 to $100 off-peak or $200 at peak - a 3× to 6× jump. Batch jobs, RAG pipelines, and agents with static system prompts are the profiles getting squeezed; one-shot summarization traffic sits closer to the 1.5–2.3× off-peak floor by comparison.
Peak hours translated to your timezone
The official page only lists UTC windows. Converted for August (daylight saving in effect):
| Zone | Peak windows (local) | Your 9-to-5 |
|---|---|---|
| Beijing (UTC+8) | 09:00–12:00, 14:00–18:00 | Peak except the 12:00–14:00 lunch |
| Central Europe (UTC+2) | 03:00–06:00, 08:00–12:00 | Mornings peak |
| London (UTC+1) | 02:00–05:00, 07:00–11:00 | Mornings peak |
| US East (UTC-4) | 21:00–00:00, 02:00–06:00 | Fully off-peak |
| US West (UTC-7) | 18:00–21:00, 23:00–03:00 | Fully off-peak |
The peak windows cover the Chinese business day minus the lunch hour (12:00–14:00 Beijing time is off-peak). US teams run their whole workday at off-peak rates, while European teams pay double on morning batches (08:00–12:00 CEST is peak) and should shift that work to the afternoon or overnight.
Is prompt caching still worth it?
Yes, unambiguously - the discount shrank, it didn't disappear. Cache-hit input was previously ~50× cheaper than cache-miss for Flash (120× for Pro); it is now ~30× cheaper for both models (31× for Flash, 30× for Pro off-peak: $0.007 vs $0.22, $0.022 vs $0.66). The audit author who measured the 2.7× jump still concluded:
"Caching is still absolutely worth it" - same r/DeepSeek audit
In their workload, caching still cut costs roughly 15× versus sending the same context uncached. Caching remains best-effort, with no guaranteed hit rate, and unused entries are generally cleared after hours to days (BenchLM) - so measure before budgeting, logging the usage object's prompt_cache_hit_tokens and prompt_cache_miss_tokens fields as BenchLM recommends to compute your actual hit share.
Is DeepSeek still the cheap option after the increase?
Off-peak, yes - at peak, the margin over GPT-5.6 narrows to nothing on the Flash tier. Output price per 1M tokens, standard rates:
| Model | Input | Output | Notes |
|---|---|---|---|
| DeepSeek V4 Flash (off-peak) | $0.22 | $0.66 | 1M context |
| DeepSeek V4 Flash (peak) | $0.44 | $1.32 | Exceeds Luna output price |
| GPT-5.6 Luna | $0.20 | $1.20 | Price cut 80% on July 30 |
| DeepSeek V4 Pro (off-peak) | $0.66 | $1.98 | Still ~6× under Terra |
| DeepSeek V4 Pro (peak) | $1.32 | $3.96 | ~3× under Terra |
| GPT-5.6 Terra | $2 | $12 | |
| Claude Sonnet 5 | $2 | $10 | Intro rate through Aug 31, then $3/$15 |
Off-peak, V4 Flash output at $0.66 undercuts GPT-5.6 Luna by 45% and Terra by 18× - DeepSeek stays cheapest on output for high-volume work. At peak, Luna ($1.20 output, $0.20 input) beats Flash on both standard rates, though Luna's cache-read rate ($0.02/M in the AI Price Index changelog) is pricier than Flash's cache-hit lane at any hour, and Flash still lists a 1M-token context and 2,500 concurrency on the official specs table. As one commenter in the announcement thread put it: "DS4 is good but when cheap."
What to do this week
- Reprice your own bill before panicking. Pull last month's usage logs and compute: cache-hit input × hit rate × hit price + cache-miss input × miss rate × miss price + output × output price, applying the peak or off-peak rate in force for each request's window. The audit's endpoints bracket the range: zero cache hits repriced at 1.5×, a 98% hit rate at 2.7×.
- Shift batchable work off peak. Cron jobs, evaluations, document processing - anything without a user waiting can run in the 17 off-peak hours. For US teams this is nearly free (your workday already is off-peak); for EU teams, move morning batches to the afternoon.
- Keep producing cache hits. Whatever prompt structure is already generating cache hits for you, keep it - the cache-hit lane still runs ~30× below cache-miss input.
- Test a Flash downgrade for routine calls. Pro runs 3× Flash across the board (3.1× on cache hits). If you haven't re-tested Flash since the 0731 post-training update, route a slice of routine traffic to it - the model-selection tradeoffs are covered in our V4 Flash vs V4 Pro comparison.
- If you leave, migration can be small. For OpenAI-compatible clients, leaving may only mean changing the base URL and model ID. Third-party hosts and relay platforms bill on their own rate cards - DataLLM Lab documented hosts quoting V4 Pro around $1.74/$3.48 per 1M even before this increase - so compare the effective rate, not the sticker.
The decision in one table:
| Your profile | Call |
|---|---|
| US timezone, cache-heavy agent traffic | Stay - off-peak rates plus the ~30× cache lane keep input cost lowest for cache-heavy traffic |
| EU timezone, morning batch jobs | Stay, reschedule - peak exposure is self-inflicted and avoidable |
| Peak-hours-only, quality-sensitive coding on Pro | Re-evaluate - Pro peak output ($3.96) is still under Terra ($12), but the gap is now 3×, not the 14× it was before August 16 |
| Zero-cache, one-shot workloads | Stay - 1.5–2.3× off-peak is the smallest increase in the table |
FAQ
When exactly did the new DeepSeek prices take effect?
August 16, 2026 at 16:00 UTC (midnight August 17 Beijing time), per both the announcement thread and the 650-call audit on r/DeepSeek. The official pricing page carried the new tables that same day.
How much will my DeepSeek bill increase?
A zero-cache workload reprices at roughly 1.5× off-peak and a measured 98%-hit-rate workload at 2.7×; peak hours double those figures, up to 12× on Pro cache-hit input. The multiplier table above has every rate pair.
Which hours are peak in US time?
US Eastern: 21:00–00:00 and 02:00–06:00. US Pacific: 18:00–21:00 and 23:00–03:00. Normal US business hours fall entirely in off-peak windows.
Are deepseek-chat and deepseek-reasoner still available?
No. The legacy aliases reached their retirement deadline on July 24, 2026, per BenchLM's model-ID history; new integrations should call deepseek-v4-flash or deepseek-v4-pro.
Is DeepSeek still cheaper than GPT-5.6 and Claude?
Off-peak, yes - Flash output is 45% under GPT-5.6 Luna and 18× under Terra. At peak, Luna's $1.20/M output undercuts Flash's $1.32/M, so the answer depends on when your traffic runs.