On identical prompts, GPT-5.6 Luna and DeepSeek V4 Pro both went 3-for-3 in our API test, so raw capability probably won't decide this matchup for you. Price structure and latency will, and DeepSeek's own pricing page now carries an official warning that a significant price increase is coming. Here is what we measured on August 10, 2026, and how to choose.
Same Prompts, Both APIs: What We Measured
We sent three identical tasks to GPT-5.6 Luna and DeepSeek V4 Pro through the same API pipeline on August 10, 2026, each model on its default settings: a Python bug fix with a deliberately terse instruction ("reply with the corrected function only"), a three-dock truck-scheduling problem with one correct answer (10:10), and a strict-JSON extraction with a nullable field. Both models got all three right.
| Task | GPT-5.6 Luna | DeepSeek V4 Pro |
|---|---|---|
Python merge_intervals bug fix | Correct (max() fix), 10.7s | Correct (max() fix), 16.1s |
| Dock-scheduling reasoning | Correct (10:10), 8.0s | Correct (10:10), 27.2s |
| Strict JSON extraction | Valid, schema-exact, 3.0s | Valid, schema-exact, 7.6s |
Two caveats and two details worth knowing. This is a three-task spot check, not a benchmark. V4 Pro runs in thinking mode by default, which explains most of its 2–3x latency gap here; you can switch it off per request, at some cost to hard reasoning. On instruction adherence, Luna returned the bare function as asked, while V4 Pro wrapped it in a markdown code fence. Harmless in a chat window, one extra parsing step in a pipeline. We re-ran the scheduling task once to check stability: both models answered 10:10 again.
Official Pricing, Verified — and DeepSeek's Coming Price Hike
Verified against both vendors' official pages on August 10, 2026: GPT-5.6 Luna costs $0.20 per million input tokens, $0.02 cached, and $1.20 output (OpenAI model page); DeepSeek V4 Pro costs $0.435 per million input tokens on cache miss, $0.003625 on cache hit, and $0.87 output (DeepSeek pricing page).
So neither model is "cheaper" across the board: Luna's uncached input is 54% cheaper; V4 Pro's output is 27% cheaper; V4 Pro's cache hits are 5.5x cheaper than Luna's. The winner depends on your token mix:
- Interactive chat (1K input + 500 output per call): Luna $0.0008 vs V4 Pro $0.00087, effectively a tie.
- Repo review (50K fresh input + 3K output): Luna $0.0136 vs V4 Pro $0.0244, Luna 44% cheaper.
- Agent loop (200K cached + 20K fresh input + 10K output per iteration): Luna $0.020 vs V4 Pro $0.018. V4 Pro pulls ahead here, and the lead grows with each additional cached iteration.
One more fact changes this math's shelf life. Footnote 2 on DeepSeek's pricing page states: "We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected." No date, no numbers. But if your case for V4 Pro rests on price, that case has an official expiration warning. Luna's pricing runs in the other direction: it arrived alongside OpenAI's August GPT-5.6 price cut.
Cache economics: where V4 Pro wins
DeepSeek V4 Pro charges $0.003625 per million cached input tokens against Luna's $0.02, a 5.5x gap that dominates costs in agent workloads, where each loop iteration re-reads a long, unchanged prefix. Developers running agents have noticed:
DeepSeek v4 Pro & MiMo v2.5 Pro ... are insanely cheap for agent-driven work due to their super low cached-input prices ($0.0036/mtok). For Luna, the cached-input price ... the pricing page puts it at $0.02/mtok, & that's 5x more expensive. — Hacker News commenter, on the GPT-5.6 launch thread
Note Luna's cache-write surcharge too: cache writes bill at 1.25x the uncached input rate (OpenAI model page), and prompts above 272K input tokens are billed at 2x input and 1.5x output for the whole request. If your agent context regularly exceeds 272K tokens, Luna's effective price roughly doubles exactly where V4 Pro's caching gets strongest.
Check prices against official docs before you model costs
While researching this comparison we found published third-party figures for these two models ranging from $0.435 to $1.74 per million input tokens for V4 Pro, and context windows for Luna listed anywhere from 400K to 1.05M. Three checks take under a minute: read pricing only from the two official pages linked above, confirm the model version string (the current API serves DeepSeek-V4-Pro; Flash is pinned at -0731), and check whether a quoted "input price" is the cache-hit or cache-miss rate; for V4 Pro those differ by 120x.
Specs That Change Workload Fit
Beyond price, four spec differences decide fit: output ceiling, modality, openness, and API surface. Both models advertise a ~1M context window (Luna: 1,050,000 tokens; V4 Pro: 1M), so context alone won't separate them.
| Spec (verified Aug 10, 2026) | GPT-5.6 Luna | DeepSeek V4 Pro |
|---|---|---|
| Context window | 1,050,000 | 1M |
| Max output tokens | 128K | 384K |
| Input modality | Text + image | Text only |
| Reasoning control | Effort levels (low→max) | Thinking on (default) / off |
| Open weights | No | Yes (self-hostable) |
| Responses API | Yes | Not yet (official: early Aug 2026) |
| Concurrency limit | Usage-tier based | 500 |
| Knowledge cutoff | Feb 16, 2026 | Not stated on pricing page |
V4 Pro's 384K output ceiling is 3x Luna's 128K. That matters for single-shot codebase generation, long translations, or bulk structured extraction where chunking adds failure modes. Luna's image input matters for a different reason, and DeepSeek users name it unprompted:
My only real complaint is that they can't do images, which limits their ability to autonomously debug some kinds of issues. — Hacker News commenter, on running DeepSeek for hobby projects
Open weights cut the other way: V4 Pro is downloadable and self-hostable, which makes it the only option here for data-residency requirements — with the caveat that serving a model this size means provisioning serious multi-GPU hardware, not a weekend project.
Benchmarks Move More With Effort Mode Than Model Choice
Published scores for GPT-5.6 Luna vs DeepSeek V4 Pro can flip winners depending on which effort mode each side was tested in — the same two models score 33.3 vs 44.3 (V4 Pro wins) on one tracker that runs Luna at low effort against V4 Pro at max, and 52 vs 45 (Luna wins) on Artificial Analysis, which runs both at max. Neither is wrong; they're measuring different configurations.
The same trap hides in latency numbers. Artificial Analysis reports Luna's time-to-first-answer-token at max effort as 132 seconds: at max effort, Luna thinks for roughly two minutes before it answers. Our test at default settings had Luna responding in 3–10.7 seconds. Both numbers are real; only one resembles your production config. Throughput favors Luna decisively once it starts emitting: 202 tokens/s vs 78 (Artificial Analysis).
Effort modes also move single-model scores enormously: V4 Pro scores 90.1% on GPQA Diamond at max thinking but 72.9% with thinking off. On shared max-effort benchmarks, Luna leads SWE-Bench Pro 62.7% to 55.4% and they tie on BrowseComp (83.3 vs 83.4). Before trusting any scorecard, ask three questions: which effort mode ran, does the latency figure include thinking time, and are the benchmark versions identical (one tracker's Terminal-Bench gap of 84.7 vs 67.9 compares v2.1 against v2.0).
Which One Should You Use?
Pick GPT-5.6 Luna for user-facing products: it was 2–3x faster on all three tasks in our test at default settings, its uncached input is half V4 Pro's price, and it accepts images. Pick DeepSeek V4 Pro for cache-heavy agent loops, outputs beyond 128K tokens, or anything that must run on your own hardware, while pricing the official increase warning into any long-term commitment. Both are one endpoint swap apart on GPT-5.6 Luna and DeepSeek V4 Pro, so you can rerun our three prompts on your own workload before committing.
| Your workload | Pick | Deciding fact |
|---|---|---|
| Chatbots, user-facing apps | Luna | 3.0–10.7s vs 7.6–27.2s latency in our test; 202 vs 78 tok/s |
| High-volume fresh-input processing | Luna | $0.20 vs $0.435 per 1M uncached input |
| Agent loops over stable context | V4 Pro | $0.003625 vs $0.02 per 1M cached input |
| Single-shot long outputs | V4 Pro | 384K vs 128K max output |
| Vision (screenshots, PDFs as images) | Luna | V4 Pro is text-only |
| Self-hosting / data residency | V4 Pro | Open weights; Luna is API-only |
| Budget certainty into 2027 | Luna | DeepSeek's official price-increase notice |
FAQ
Which is cheaper, GPT-5.6 Luna or DeepSeek V4 Pro?
Neither, universally. Luna's uncached input is 54% cheaper ($0.20 vs $0.435 per 1M tokens); V4 Pro's cached input is 5.5x cheaper ($0.003625 vs $0.02) and its output is 27% cheaper ($0.87 vs $1.20). Cache-heavy agents run cheaper on V4 Pro; fresh-input workloads run cheaper on Luna.
Do both models support a 1M-token context window?
Yes: Luna offers 1,050,000 tokens and V4 Pro offers 1M. Note Luna requests exceeding 272K input tokens are billed at 2x input and 1.5x output for the entire request.
Is DeepSeek V4 Pro open source, and can I self-host it?
Yes. V4 Pro ships open weights and can be self-hosted, unlike API-only Luna, but plan for multi-GPU serving hardware rather than a single workstation.
Does GPT-5.6 Luna support image input?
Yes, Luna accepts text and image input. DeepSeek V4 Pro is text-only. Its own users cite the missing vision support as the main limitation for autonomous debugging workflows.
What are the maximum output tokens for each model?
DeepSeek V4 Pro can emit up to 384K tokens per response; GPT-5.6 Luna caps at 128K. For single-shot generation of long documents or large code files, V4 Pro's ceiling is 3x higher.
Related reading
- DeepSeek V4 Pro vs GPT-5.6 Sol: the same DeepSeek flagship against OpenAI's mid tier
- GPT-5.6 Luna review: Luna on its own terms, beyond this matchup
- DeepSeek V4 Flash vs V4 Pro: if V4 Pro's pricing risk pushes you down-tier