AIREITER

GPT-5.6 Luna vs DeepSeek V4 Pro: We Tested Both APIs

Last Updated: 2026-08-10 10:48:31

On identical prompts, GPT-5.6 Luna and DeepSeek V4 Pro both went 3-for-3 in our API test, so raw capability probably won't decide this matchup for you. Price structure and latency will, and DeepSeek's own pricing page now carries an official warning that a significant price increase is coming. Here is what we measured on August 10, 2026, and how to choose.

Same Prompts, Both APIs: What We Measured

We sent three identical tasks to GPT-5.6 Luna and DeepSeek V4 Pro through the same API pipeline on August 10, 2026, each model on its default settings: a Python bug fix with a deliberately terse instruction ("reply with the corrected function only"), a three-dock truck-scheduling problem with one correct answer (10:10), and a strict-JSON extraction with a nullable field. Both models got all three right.

TaskGPT-5.6 LunaDeepSeek V4 Pro
Python merge_intervals bug fixCorrect (max() fix), 10.7sCorrect (max() fix), 16.1s
Dock-scheduling reasoningCorrect (10:10), 8.0sCorrect (10:10), 27.2s
Strict JSON extractionValid, schema-exact, 3.0sValid, schema-exact, 7.6s
Latency comparison across three identical prompts, GPT-5.6 Luna vs DeepSeek V4 Pro

Two caveats and two details worth knowing. This is a three-task spot check, not a benchmark. V4 Pro runs in thinking mode by default, which explains most of its 2–3x latency gap here; you can switch it off per request, at some cost to hard reasoning. On instruction adherence, Luna returned the bare function as asked, while V4 Pro wrapped it in a markdown code fence. Harmless in a chat window, one extra parsing step in a pipeline. We re-ran the scheduling task once to check stability: both models answered 10:10 again.

Official Pricing, Verified — and DeepSeek's Coming Price Hike

Verified against both vendors' official pages on August 10, 2026: GPT-5.6 Luna costs $0.20 per million input tokens, $0.02 cached, and $1.20 output (OpenAI model page); DeepSeek V4 Pro costs $0.435 per million input tokens on cache miss, $0.003625 on cache hit, and $0.87 output (DeepSeek pricing page).

Official API pricing per million tokens, GPT-5.6 Luna vs DeepSeek V4 Pro

So neither model is "cheaper" across the board: Luna's uncached input is 54% cheaper; V4 Pro's output is 27% cheaper; V4 Pro's cache hits are 5.5x cheaper than Luna's. The winner depends on your token mix:

  • Interactive chat (1K input + 500 output per call): Luna $0.0008 vs V4 Pro $0.00087, effectively a tie.
  • Repo review (50K fresh input + 3K output): Luna $0.0136 vs V4 Pro $0.0244, Luna 44% cheaper.
  • Agent loop (200K cached + 20K fresh input + 10K output per iteration): Luna $0.020 vs V4 Pro $0.018. V4 Pro pulls ahead here, and the lead grows with each additional cached iteration.
GPT-5.6 Luna official model page with pricing and context window

One more fact changes this math's shelf life. Footnote 2 on DeepSeek's pricing page states: "We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected." No date, no numbers. But if your case for V4 Pro rests on price, that case has an official expiration warning. Luna's pricing runs in the other direction: it arrived alongside OpenAI's August GPT-5.6 price cut.

DeepSeek V4 Pro official pricing table with the price-increase footnote

Cache economics: where V4 Pro wins

DeepSeek V4 Pro charges $0.003625 per million cached input tokens against Luna's $0.02, a 5.5x gap that dominates costs in agent workloads, where each loop iteration re-reads a long, unchanged prefix. Developers running agents have noticed:

DeepSeek v4 Pro & MiMo v2.5 Pro ... are insanely cheap for agent-driven work due to their super low cached-input prices ($0.0036/mtok). For Luna, the cached-input price ... the pricing page puts it at $0.02/mtok, & that's 5x more expensive. — Hacker News commenter, on the GPT-5.6 launch thread

Note Luna's cache-write surcharge too: cache writes bill at 1.25x the uncached input rate (OpenAI model page), and prompts above 272K input tokens are billed at 2x input and 1.5x output for the whole request. If your agent context regularly exceeds 272K tokens, Luna's effective price roughly doubles exactly where V4 Pro's caching gets strongest.

Check prices against official docs before you model costs

While researching this comparison we found published third-party figures for these two models ranging from $0.435 to $1.74 per million input tokens for V4 Pro, and context windows for Luna listed anywhere from 400K to 1.05M. Three checks take under a minute: read pricing only from the two official pages linked above, confirm the model version string (the current API serves DeepSeek-V4-Pro; Flash is pinned at -0731), and check whether a quoted "input price" is the cache-hit or cache-miss rate; for V4 Pro those differ by 120x.

Specs That Change Workload Fit

Beyond price, four spec differences decide fit: output ceiling, modality, openness, and API surface. Both models advertise a ~1M context window (Luna: 1,050,000 tokens; V4 Pro: 1M), so context alone won't separate them.

Spec (verified Aug 10, 2026)GPT-5.6 LunaDeepSeek V4 Pro
Context window1,050,0001M
Max output tokens128K384K
Input modalityText + imageText only
Reasoning controlEffort levels (low→max)Thinking on (default) / off
Open weightsNoYes (self-hostable)
Responses APIYesNot yet (official: early Aug 2026)
Concurrency limitUsage-tier based500
Knowledge cutoffFeb 16, 2026Not stated on pricing page

V4 Pro's 384K output ceiling is 3x Luna's 128K. That matters for single-shot codebase generation, long translations, or bulk structured extraction where chunking adds failure modes. Luna's image input matters for a different reason, and DeepSeek users name it unprompted:

My only real complaint is that they can't do images, which limits their ability to autonomously debug some kinds of issues. — Hacker News commenter, on running DeepSeek for hobby projects

Open weights cut the other way: V4 Pro is downloadable and self-hostable, which makes it the only option here for data-residency requirements — with the caveat that serving a model this size means provisioning serious multi-GPU hardware, not a weekend project.

Benchmarks Move More With Effort Mode Than Model Choice

Published scores for GPT-5.6 Luna vs DeepSeek V4 Pro can flip winners depending on which effort mode each side was tested in — the same two models score 33.3 vs 44.3 (V4 Pro wins) on one tracker that runs Luna at low effort against V4 Pro at max, and 52 vs 45 (Luna wins) on Artificial Analysis, which runs both at max. Neither is wrong; they're measuring different configurations.

The same trap hides in latency numbers. Artificial Analysis reports Luna's time-to-first-answer-token at max effort as 132 seconds: at max effort, Luna thinks for roughly two minutes before it answers. Our test at default settings had Luna responding in 3–10.7 seconds. Both numbers are real; only one resembles your production config. Throughput favors Luna decisively once it starts emitting: 202 tokens/s vs 78 (Artificial Analysis).

Effort modes also move single-model scores enormously: V4 Pro scores 90.1% on GPQA Diamond at max thinking but 72.9% with thinking off. On shared max-effort benchmarks, Luna leads SWE-Bench Pro 62.7% to 55.4% and they tie on BrowseComp (83.3 vs 83.4). Before trusting any scorecard, ask three questions: which effort mode ran, does the latency figure include thinking time, and are the benchmark versions identical (one tracker's Terminal-Bench gap of 84.7 vs 67.9 compares v2.1 against v2.0).

Which One Should You Use?

Pick GPT-5.6 Luna for user-facing products: it was 2–3x faster on all three tasks in our test at default settings, its uncached input is half V4 Pro's price, and it accepts images. Pick DeepSeek V4 Pro for cache-heavy agent loops, outputs beyond 128K tokens, or anything that must run on your own hardware, while pricing the official increase warning into any long-term commitment. Both are one endpoint swap apart on GPT-5.6 Luna and DeepSeek V4 Pro, so you can rerun our three prompts on your own workload before committing.

Your workloadPickDeciding fact
Chatbots, user-facing appsLuna3.0–10.7s vs 7.6–27.2s latency in our test; 202 vs 78 tok/s
High-volume fresh-input processingLuna$0.20 vs $0.435 per 1M uncached input
Agent loops over stable contextV4 Pro$0.003625 vs $0.02 per 1M cached input
Single-shot long outputsV4 Pro384K vs 128K max output
Vision (screenshots, PDFs as images)LunaV4 Pro is text-only
Self-hosting / data residencyV4 ProOpen weights; Luna is API-only
Budget certainty into 2027LunaDeepSeek's official price-increase notice

FAQ

Which is cheaper, GPT-5.6 Luna or DeepSeek V4 Pro?

Neither, universally. Luna's uncached input is 54% cheaper ($0.20 vs $0.435 per 1M tokens); V4 Pro's cached input is 5.5x cheaper ($0.003625 vs $0.02) and its output is 27% cheaper ($0.87 vs $1.20). Cache-heavy agents run cheaper on V4 Pro; fresh-input workloads run cheaper on Luna.

Do both models support a 1M-token context window?

Yes: Luna offers 1,050,000 tokens and V4 Pro offers 1M. Note Luna requests exceeding 272K input tokens are billed at 2x input and 1.5x output for the entire request.

Is DeepSeek V4 Pro open source, and can I self-host it?

Yes. V4 Pro ships open weights and can be self-hosted, unlike API-only Luna, but plan for multi-GPU serving hardware rather than a single workstation.

Does GPT-5.6 Luna support image input?

Yes, Luna accepts text and image input. DeepSeek V4 Pro is text-only. Its own users cite the missing vision support as the main limitation for autonomous debugging workflows.

What are the maximum output tokens for each model?

DeepSeek V4 Pro can emit up to 384K tokens per response; GPT-5.6 Luna caps at 128K. For single-shot generation of long documents or large code files, V4 Pro's ceiling is 3x higher.

Related reading