Pick DeepSeek V4 Pro if the bill matters and Grok if latency does, a split that held across seven tasks I ran through both APIs. The caveat is that DeepSeek's own pricing page warns of a significant increase, which would close most of the gap the recommendation rests on.
What You're Actually Comparing in August 2026
Both models changed identity in the week before this comparison. Grok 4.6 now occupies the Latest slot in xAI's model documentation, described as the flagship "for code and everything else," and DeepSeek V4 Pro is currently served as the DeepSeek-V4-Pro-0813 snapshot. Neither change shows up in an API model ID.
The DeepSeek version string lives only in the MODEL VERSION row of the pricing page. The API still answers to a plain deepseek-v4-pro with no date suffix, so a request you send today and a request you sent in July both look identical in your logs while potentially hitting different weights.
Here is the full specification and price comparison, taken from xAI's pricing page and DeepSeek's pricing page on 13 August 2026:
| Grok 4.6 | DeepSeek V4 Pro (0813) | |
|---|---|---|
| Context window | 500K | 1M |
| Max output | Not published | 384K |
| Input / 1M | $2.00 | $0.435 |
| Cached input / 1M | $0.50 | $0.003625 |
| Output / 1M | $6.00 | $0.87 |
| Long-context rate | 2x all rates at ≥200K prompt | None |
| Image input | Yes | No |
| Tool calls | Yes | Yes |
| Reasoning | Configurable | Thinking / non-thinking |
| Knowledge cutoff | 1 February 2026 | Not published |
Grok 4.6 kept Grok 4.5's headline rates of $2.00 and $6.00 unchanged. The one price that moved is cached input: the same xAI table lists $0.30 for Grok 4.5 and $0.50 for Grok 4.6, making cache reuse 67% more expensive on the newer model.
Third-party gateways are still serving Grok 4.5
Availability in xAI's docs does not mean availability wherever you buy tokens. On the gateway I tested through, grok-4.6 returned a model_not_found error, and the grok-latest and grok aliases both resolved to grok-4.5-build. If you reach Grok through a reseller rather than xAI's own API, check what your alias actually resolves to before assuming a version.
That lag is also why the measurements below ran against grok-4.5 rather than 4.6. Since both versions carry identical input and output pricing, the cost conclusions transfer; the per-task accuracy results are 4.5's.
Instruction Adherence Is Where They Actually Split
Across seven mechanically-checked tasks, six came out even and one split decisively. The split task asked for exactly four lines, each ending with a description of exactly five words. DeepSeek V4 Pro passed 5 of 5 attempts; Grok passed 0 of 5.
Grok's failures were not one bug repeating. Two runs emitted a raw web_search tool call as body text instead of an answer. One of them read, in full, "I need accurate information on the four stages of canary deployment. Let me search for that." Two runs produced the right four-line shape but miscounted the descriptions at four and six words. One run went to 130 seconds and 10,357 output tokens, returning 331 lines.
The leakage is probably the gateway's doing rather than the model's: it inflates a 6-token prompt to 201 reported input tokens, which points to an injected preamble declaring xAI's server-side tools. The word-count misses are separate. Those runs answered normally and still dropped the constraint.
Method: deepseek-v4-pro and grok-4.5 through the same OpenAI-compatible gateway at default settings, one run per task except the format task, which ran five times each. Every task had a mechanically checked answer, so scoring involved no judge model. The six even tasks appear by name in the cost table below.
Both models accepted a false premise
The seventh task planted two errors in a prompt: that DeepSeek V4 Pro has a 128K context window, and that Grok 4.6 launched in January 2026 with 2M context. Then I asked which to use for a 900,000-token corpus. Neither model challenged either figure. Both reasoned cleanly from the bad numbers to a confident recommendation of Grok.
If you paste stale specifications into a prompt and ask for a recommendation, both models will hand you a well-argued wrong answer rather than correcting the input.
Grok Used Fewer Output Tokens and Still Cost 5.7x More
Across the six tasks both models completed, Grok produced 5,275 output tokens to DeepSeek's 6,331, or 17% fewer. Priced at each model's official output rate, those same responses cost $0.0317 on Grok and $0.0055 on DeepSeek, a 5.7x gap the other way.
| Task | DeepSeek tokens | Grok tokens | DeepSeek cost | Grok cost |
|---|---|---|---|---|
| Strict JSON | 147 | 219 | 0.013¢ | 0.131¢ |
| Merge ranges | 384 | 416 | 0.033¢ | 0.250¢ |
| 56K needle | 181 | 174 | 0.016¢ | 0.104¢ |
| Pricing arithmetic | 520 | 647 | 0.045¢ | 0.388¢ |
| Duration parser | 4,087 | 2,611 | 0.356¢ | 1.567¢ |
| SQL aggregate | 1,012 | 1,208 | 0.088¢ | 0.725¢ |
| Total | 6,331 | 5,275 | 0.551¢ | 3.165¢ |
Token efficiency is real and it is not enough: Grok won the duration parser on concision, 2,611 tokens against 4,087, and still cost four times as much for that answer.
Scale it to a working month and the arithmetic is easy to redo: 1,000 requests at 50,000 uncached input tokens and 3,000 output tokens each is 50M input and 3M output. On Grok 4.6 that is $100.00 plus $18.00, or $118.00. On DeepSeek V4 Pro it is $21.75 plus $2.61, or $24.36.
Grok was faster on all six tasks, and the margin widened with task length: 48.8 seconds against 83.3 on the duration parser, versus 3.7 against 5.4 on the JSON task. Single runs on a shared gateway, so treat the small gaps as noise and the large one as directional.
Two Billing Rules the Spec Tables Don't Show
Two pricing mechanics decide more of the real bill than the headline rates do, and neither appears as a column in a specification table.
xAI's long-context threshold is a cliff, not a slope. Any request whose prompt reaches 200,000 tokens is billed at double rates for every token in that request, not just the tokens past the line. A 199,000-token prompt costs $0.398 in input; a 201,000-token prompt costs $0.804. Crossing by 1% doubles the bill.
DeepSeek's cache is nearly free. A cache hit costs $0.003625 per 1M input tokens against Grok 4.6's $0.50, about 138 times cheaper. On repeated work against a stable codebase, that is the number that decides the invoice. One developer on Hacker News reported:
DeepSeek's official API has a cache hit rate of over 99% if you use it continuously within the same codebase for long sessions, so it's much cheaper than frontier models. I have an example of 200M token session in claude code.
The catch is that DeepSeek's rates are explicitly temporary. The pricing page carries a footnote: "We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected." A Hacker News breakdown of DeepSeek's promotional pricing puts the pre-discount rates at $1.74 input and $3.48 output. The arithmetic is self-consistent, since $0.435 and $0.87 are exactly a quarter of those figures, but DeepSeek has published no promotion terms.
If the increase lands at those levels, the monthly scenario above moves from $24.36 to $97.44 against Grok's $118.00. The output-price gap narrows from 6.9x to about 1.7x, and a decision made purely on price stops being obvious.
Which One to Use
Use DeepSeek V4 Pro for high-volume, cost-sensitive, repeat-context work: batch processing, long agentic coding sessions against one repository, anything above roughly 200K tokens per prompt where Grok's threshold doubles the rate. The 1M context window and the near-free cache compound in your favour, and it held the exact-format constraint in all five attempts. You can run it directly from the DeepSeek V4 Pro API page.
Use Grok 4.6 for latency-sensitive and multimodal work. Grok was faster on every task, though that and the format result were measured on 4.5, not 4.6. It accepts image input where DeepSeek does not, and its published knowledge cutoff of 1 February 2026 is at least stated. If you are routing through a gateway, confirm you are getting 4.6 and not 4.5. Grok on AIReiter and every other reseller inherits whatever version is wired up upstream.
Do not choose either on the strength of an announced price. DeepSeek's own footnote makes today's advantage a moving target, and Grok's cached-input rate already rose 67% between 4.5 and 4.6. Re-run the arithmetic above against the live pricing pages before committing a workload.
FAQ
Is Grok 4.6 actually released?
Yes. As of 13 August 2026 Grok 4.6 appears in xAI's model documentation in the Latest position with published pricing of $2.00 input and $6.00 output per 1M tokens and a 500K context window. Third-party gateways may still route grok-latest to Grok 4.5.
What is DeepSeek-V4-Pro-0813?
It is the model version string DeepSeek currently serves behind the deepseek-v4-pro API endpoint, listed in the MODEL VERSION row of the official pricing page. The API ID carries no date suffix, so the snapshot can change without any change to your request.
Which model has the larger context window?
DeepSeek V4 Pro, at 1M tokens against Grok 4.6's 500K. DeepSeek also publishes a 384K maximum output; xAI does not publish a maximum output figure for Grok 4.6.
Is DeepSeek V4 Pro cheaper than Grok 4.6?
At list prices on 13 August 2026, yes: 4.6x cheaper on input and 6.9x on output. DeepSeek's pricing page warns of a significant increase, and at the pre-promotion rates developers have identified the output gap would fall to roughly 1.7x.
Which one follows instructions more reliably?
In the one test that separated them, DeepSeek V4 Pro held an exact-format constraint in 5 of 5 attempts while Grok held it in 0 of 5. Both models failed to challenge false specifications planted in a prompt.
Related reading: GPT-5.6 Luna vs DeepSeek V4 Pro