Grok 4.6 landed on 12 August 2026, five days after the date Elon Musk's "two weeks" implied, and the price sheet looks unchanged: $2.00 input, $6.00 output per million tokens, 500K context, exactly Grok 4.5's terms. The bill is not unchanged. Cached input costs 67% more than on 4.5, and independent measurement shows the model spending more tokens to finish the same work.
What shipped on 12 August
The model ID is grok-4.6, and SpaceXAI's release notes date its API availability to 12 August 2026. The documented specification:
| Property | Grok 4.6 |
|---|---|
| Model ID | grok-4.6 |
| Context | 500K tokens |
| Input | Text and image |
| Output | Text only, no stated output limit |
| Reasoning effort | low, medium, high (default), xhigh |
| Knowledge cutoff | 1 February 2026 |
Two of those rows are new. xhigh is a fourth reasoning tier that Grok 4.5 does not offer — 4.5 tops out at high. And the documentation now carries a knowledge cutoff of 1 February 2026, which matters if you were routing dated questions around the model's blind spot.
The model catalog lists Grok 4.6 as the flagship for code and chat, tagged Latest in the sidebar, with the page stamped 12 August 2026. Grok 4.5 has not been retired; it still appears on the pricing page at its own rates.
Distribution went wide on day one. SpaceXAI's announcement lists Cursor and Grok Build, with 2x included usage in both for the first week, plus the API and partners including OpenRouter, Vercel, and Cloudflare. It also mentions a fast variant at twice the price; that variant has no row on the public pricing page yet, so treat the 2x figure as the announcement's number rather than a published rate.
Grok 4.6 pricing: same headline, pricier cache
Input and output rates carry over from Grok 4.5 unchanged. The cached-input rate does not.
| Rate per 1M tokens | Grok 4.5 | Grok 4.6 |
|---|---|---|
| Input (under 200K) | $2.00 | $2.00 |
| Cached input (under 200K) | $0.30 | $0.50 |
| Output (under 200K) | $6.00 | $6.00 |
| Input / cached / output (200K+) | $4.00 / $0.60 / $12.00 | $4.00 / $1.00 / $12.00 |
Cached input went from an 85% discount to a 75% discount: 20 cents more per million tokens, on the line item that scales with agent steps rather than with tasks. A harness replaying 50M cached tokens a day pays $15.00 on 4.6 against $9.00 on 4.5.
The long-context rule is unchanged and still catches people. Once a request's prompt reaches 200K tokens, every token in that request bills at the higher tier — a 210K-token call is $4.00 per million input across the whole request, not a blend of the two tiers. OpenRouter's public model listing reports the same rates and the same 200K override, so a third-party route does not change the arithmetic.
Where Grok 4.6 wins and loses on the official evals
SpaceXAI published a ten-row comparison against Grok 4.5, GPT-5.6 Sol Max, and Fable 5 Max. Grok 4.6 takes the top score in three of them: GDPVal-AA v2 (1753), AA-Briefcase (1577), and Harvey LAB (15.8%). It loses the other seven. The eight rows that bear on routing:
| Eval | Grok 4.6 | Grok 4.5 | GPT-5.6 Sol | Fable 5 |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
Terminal-Bench is the gap worth reading twice. Grok 4.6 scores 26% where GPT-5.6 Sol scores 34.6% and Fable 5 scores 34.1% — a third-place finish by eight points on the eval that most resembles an agent driving a shell. DeepSWE tells the same story with different numbers: 65.9% against Sol's 73%. If your workload is terminal-driven agents, the vendor's own table says this is not the model that leads.
The gain over Grok 4.5 is real and consistent, though. Every row improves, and APEX-Agents (47.1% to 57.5%) and Terminal-Bench (15.7% to 26%) improve by more than ten points. SpaceXAI attributes this to a longer supplemental training run, regenerated SFT trajectories, and agentic RL across coding and knowledge-work environments. Note what the announcement never states: a parameter count. The 2-trillion figure attached to this model came from Musk, not from the release.
The cost of that intelligence, measured
Vendor tables report scores. They do not report what the score cost. Artificial Analysis publishes both, and its 13 August reading is where Grok 4.6's real trade-off shows up.
| Measured on the same index | Grok 4.5 (high) | Grok 4.6 (high) |
|---|---|---|
| Intelligence Index | 56 (#16 of 184) | 61 (#6 of 184) |
| Output speed | 56.9 tokens/s (#102) | 67.6 tokens/s (#74) |
| Output tokens generated | 60M | 72M |
| Cost of the full evaluation | $579.21 | $1,068.47 |
Five points of intelligence cost 1.8x. Identical headline rates, and the same benchmark suite billed $489 more: partly verbosity (72M output tokens against 60M), partly the cache repricing above. Artificial Analysis publishes no line-item breakdown of the run, so 1.8x is an observed total, not a formula you can rederive.
Speed moved in your favour: 67.6 output tokens per second against 56.9 for 4.5, which lifts it from the bottom half to roughly the median of the 184 models in its class.
What Musk said, against what shipped
The name entered the record as a three-word reply. On 18 July, replying to @minchoi, Musk wrote: "Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5)." Asked by Andrew Curran whether that was Grok 4.6, he answered: "Yeah, Grok 4.6."
On 24 July came the schedule: "Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks." Two weeks from that date lands near 7 August; the model shipped on 12 August, five days later.
Three claims from those posts did not survive the release intact:
- 2 trillion parameters is still unconfirmed. SpaceXAI's announcement, model documentation, and release notes give no parameter count for Grok 4.6.
- "Token efficiency close to Grok 4.5" is contradicted by Artificial Analysis's measurement: 72M output tokens against 60M on the same evaluation.
- "Might exceed Kimi" remains untested against Moonshot's K3 in SpaceXAI's table, which compares only against GPT-5.6 Sol and Fable 5.
Should you move from Grok 4.5 to Grok 4.6
Yes for agentic and knowledge work if you are coming from Grok 4.5, with the token bill checked after a week rather than at migration. Grok 4.5 remains available at its old rates, so the switch is reversible and worth measuring rather than assuming.
Move now if your work is long-horizon agents or knowledge tasks: APEX-Agents up 10.4 points, GDPVal-AA up 227, and the faster generation rate compound over multi-step runs. Stay on 4.5 if your traffic is cache-heavy and price-sensitive — the 67% cache-read increase is a genuine regression, and 4.5's scores did not drop when 4.6 shipped. Look elsewhere for terminal agents specifically, where SpaceXAI's own table puts Grok 4.6 eight points behind GPT-5.6 Sol and Fable 5.
Three things to check before you flip a production route:
- Know that both IDs are aliases. SpaceXAI's alias rules point
grok-4.6at the latest stable build andgrok-4.6-latestat the newest one; only a dated<modelname>-<date>ID is frozen, and neither the model catalog nor OpenRouter lists a dated 4.6 variant. Reproducible workflows have no pin available today. - Re-test your reasoning effort. The new
xhightier is not a free upgrade — it spends more tokens on a model already measured as more verbose than its predecessor. Defaulthighis the like-for-like comparison against 4.5. - Verify your provider has it. Day-one availability at SpaceXAI is not day-one availability everywhere. Querying
/v1/modelson 13 August, the official API and OpenRouter both returnedgrok-4.6, while an OpenAI-compatible reseller endpoint I tested still topped out atgrok-4.5and rejected the new ID withmodel_not_found. Confirm the exact ID answers on your own key before routing traffic.
FAQ
How much does the Grok 4.6 API cost?
$2.00 per million input tokens and $6.00 per million output tokens under 200K context, with cached input at $0.50. Above 200K prompt tokens, the whole request bills at $4.00 input, $1.00 cached, and $12.00 output. A fast variant at twice the price is mentioned in the announcement but is not on the public pricing page.
Is Grok 4.6 better than Grok 4.5?
On every benchmark SpaceXAI published, yes: 61 against 56 on the AA Intelligence Index, with the largest gains on agentic evals. Independent measurement adds the qualifier — it generated 72M output tokens against Grok 4.5's 60M on the same evaluation, so the same task costs more.
How many parameters does Grok 4.6 have?
Undisclosed. Musk described a 2-trillion-parameter model on 18 July 2026 and confirmed the name, but SpaceXAI's release materials give no parameter count, no architecture details, and no statement on how many parameters activate per token.
What is Grok 4.6's context window and knowledge cutoff?
500K tokens of context, with text and image input and text-only output. The documentation gives a knowledge cutoff of 1 February 2026, and the model has no access to events after that without server-side Web Search or X Search enabled.
Is Grok made by xAI or SpaceXAI?
Both names refer to the same company. The developer documentation is titled "SpaceXAI Docs" and the announcement footer still reads X.AI LLC, after the 2026 move under SpaceX. The x.ai domains and the grok-* model IDs were not renamed.
When is Grok 4.7 coming out?
No date has been announced. Musk's 24 July post put Grok 4.7 about four weeks out, pointing at around 21 August 2026, and Grok 4.6 arrived five days past its own implied date.
Related reading
Pricing, model IDs and benchmark figures verified against SpaceXAI documentation, the Grok 4.6 announcement and Artificial Analysis on 13 August 2026.