AIREITER

Grok 4.6 API: Pricing, Benchmarks and What Changed

Last Updated: 2026-08-13 02:25:10

Grok 4.6 landed on 12 August 2026, five days after the date Elon Musk's "two weeks" implied, and the price sheet looks unchanged: $2.00 input, $6.00 output per million tokens, 500K context, exactly Grok 4.5's terms. The bill is not unchanged. Cached input costs 67% more than on 4.5, and independent measurement shows the model spending more tokens to finish the same work.

What shipped on 12 August

The model ID is grok-4.6, and SpaceXAI's release notes date its API availability to 12 August 2026. The documented specification:

PropertyGrok 4.6
Model IDgrok-4.6
Context500K tokens
InputText and image
OutputText only, no stated output limit
Reasoning effortlow, medium, high (default), xhigh
Knowledge cutoff1 February 2026
SpaceXAI release notes entry dated August 12, 2026 announcing Grok 4.6 on the xAI API with its context window and pricing

Two of those rows are new. xhigh is a fourth reasoning tier that Grok 4.5 does not offer — 4.5 tops out at high. And the documentation now carries a knowledge cutoff of 1 February 2026, which matters if you were routing dated questions around the model's blind spot.

The model catalog lists Grok 4.6 as the flagship for code and chat, tagged Latest in the sidebar, with the page stamped 12 August 2026. Grok 4.5 has not been retired; it still appears on the pricing page at its own rates.

SpaceXAI developer documentation showing Grok 4.6 tagged as New with 500k context, $2.00 input and $6.00 output pricing

Distribution went wide on day one. SpaceXAI's announcement lists Cursor and Grok Build, with 2x included usage in both for the first week, plus the API and partners including OpenRouter, Vercel, and Cloudflare. It also mentions a fast variant at twice the price; that variant has no row on the public pricing page yet, so treat the 2x figure as the announcement's number rather than a published rate.

Grok 4.6 pricing: same headline, pricier cache

Input and output rates carry over from Grok 4.5 unchanged. The cached-input rate does not.

Rate per 1M tokensGrok 4.5Grok 4.6
Input (under 200K)$2.00$2.00
Cached input (under 200K)$0.30$0.50
Output (under 200K)$6.00$6.00
Input / cached / output (200K+)$4.00 / $0.60 / $12.00$4.00 / $1.00 / $12.00
Grouped bar chart comparing Grok 4.5 and Grok 4.6 input, cached input and output rates under 200K context

Cached input went from an 85% discount to a 75% discount: 20 cents more per million tokens, on the line item that scales with agent steps rather than with tasks. A harness replaying 50M cached tokens a day pays $15.00 on 4.6 against $9.00 on 4.5.

The long-context rule is unchanged and still catches people. Once a request's prompt reaches 200K tokens, every token in that request bills at the higher tier — a 210K-token call is $4.00 per million input across the whole request, not a blend of the two tiers. OpenRouter's public model listing reports the same rates and the same 200K override, so a third-party route does not change the arithmetic.

Where Grok 4.6 wins and loses on the official evals

SpaceXAI published a ten-row comparison against Grok 4.5, GPT-5.6 Sol Max, and Fable 5 Max. Grok 4.6 takes the top score in three of them: GDPVal-AA v2 (1753), AA-Briefcase (1577), and Harvey LAB (15.8%). It loses the other seven. The eight rows that bear on routing:

Official SpaceXAI evaluation table comparing Grok 4.6 with Grok 4.5, GPT-5.6 Sol and Fable 5 across ten benchmarks
EvalGrok 4.6Grok 4.5GPT-5.6 SolFable 5
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.161.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
AA-Briefcase1577131315021574

Terminal-Bench is the gap worth reading twice. Grok 4.6 scores 26% where GPT-5.6 Sol scores 34.6% and Fable 5 scores 34.1% — a third-place finish by eight points on the eval that most resembles an agent driving a shell. DeepSWE tells the same story with different numbers: 65.9% against Sol's 73%. If your workload is terminal-driven agents, the vendor's own table says this is not the model that leads.

The gain over Grok 4.5 is real and consistent, though. Every row improves, and APEX-Agents (47.1% to 57.5%) and Terminal-Bench (15.7% to 26%) improve by more than ten points. SpaceXAI attributes this to a longer supplemental training run, regenerated SFT trajectories, and agentic RL across coding and knowledge-work environments. Note what the announcement never states: a parameter count. The 2-trillion figure attached to this model came from Musk, not from the release.

The cost of that intelligence, measured

Vendor tables report scores. They do not report what the score cost. Artificial Analysis publishes both, and its 13 August reading is where Grok 4.6's real trade-off shows up.

Measured on the same indexGrok 4.5 (high)Grok 4.6 (high)
Intelligence Index56 (#16 of 184)61 (#6 of 184)
Output speed56.9 tokens/s (#102)67.6 tokens/s (#74)
Output tokens generated60M72M
Cost of the full evaluation$579.21$1,068.47
Bar chart comparing the cost of running the Artificial Analysis Intelligence Index on Grok 4.5 and Grok 4.6

Five points of intelligence cost 1.8x. Identical headline rates, and the same benchmark suite billed $489 more: partly verbosity (72M output tokens against 60M), partly the cache repricing above. Artificial Analysis publishes no line-item breakdown of the run, so 1.8x is an observed total, not a formula you can rederive.

Speed moved in your favour: 67.6 output tokens per second against 56.9 for 4.5, which lifts it from the bottom half to roughly the median of the 184 models in its class.

What Musk said, against what shipped

The name entered the record as a three-word reply. On 18 July, replying to @minchoi, Musk wrote: "Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5)." Asked by Andrew Curran whether that was Grok 4.6, he answered: "Yeah, Grok 4.6."

Elon Musk's X thread describing a 2T model better than the 1.5T Grok 4.5, followed by his three-word reply confirming the name Grok 4.6

On 24 July came the schedule: "Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks." Two weeks from that date lands near 7 August; the model shipped on 12 August, five days later.

Elon Musk's X post stating Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks, replying to his FrontierCode Pareto frontier chart

Three claims from those posts did not survive the release intact:

  • 2 trillion parameters is still unconfirmed. SpaceXAI's announcement, model documentation, and release notes give no parameter count for Grok 4.6.
  • "Token efficiency close to Grok 4.5" is contradicted by Artificial Analysis's measurement: 72M output tokens against 60M on the same evaluation.
  • "Might exceed Kimi" remains untested against Moonshot's K3 in SpaceXAI's table, which compares only against GPT-5.6 Sol and Fable 5.

Should you move from Grok 4.5 to Grok 4.6

Yes for agentic and knowledge work if you are coming from Grok 4.5, with the token bill checked after a week rather than at migration. Grok 4.5 remains available at its old rates, so the switch is reversible and worth measuring rather than assuming.

Move now if your work is long-horizon agents or knowledge tasks: APEX-Agents up 10.4 points, GDPVal-AA up 227, and the faster generation rate compound over multi-step runs. Stay on 4.5 if your traffic is cache-heavy and price-sensitive — the 67% cache-read increase is a genuine regression, and 4.5's scores did not drop when 4.6 shipped. Look elsewhere for terminal agents specifically, where SpaceXAI's own table puts Grok 4.6 eight points behind GPT-5.6 Sol and Fable 5.

Three things to check before you flip a production route:

  1. Know that both IDs are aliases. SpaceXAI's alias rules point grok-4.6 at the latest stable build and grok-4.6-latest at the newest one; only a dated <modelname>-<date> ID is frozen, and neither the model catalog nor OpenRouter lists a dated 4.6 variant. Reproducible workflows have no pin available today.
  2. Re-test your reasoning effort. The new xhigh tier is not a free upgrade — it spends more tokens on a model already measured as more verbose than its predecessor. Default high is the like-for-like comparison against 4.5.
  3. Verify your provider has it. Day-one availability at SpaceXAI is not day-one availability everywhere. Querying /v1/models on 13 August, the official API and OpenRouter both returned grok-4.6, while an OpenAI-compatible reseller endpoint I tested still topped out at grok-4.5 and rejected the new ID with model_not_found. Confirm the exact ID answers on your own key before routing traffic.

FAQ

How much does the Grok 4.6 API cost?

$2.00 per million input tokens and $6.00 per million output tokens under 200K context, with cached input at $0.50. Above 200K prompt tokens, the whole request bills at $4.00 input, $1.00 cached, and $12.00 output. A fast variant at twice the price is mentioned in the announcement but is not on the public pricing page.

Is Grok 4.6 better than Grok 4.5?

On every benchmark SpaceXAI published, yes: 61 against 56 on the AA Intelligence Index, with the largest gains on agentic evals. Independent measurement adds the qualifier — it generated 72M output tokens against Grok 4.5's 60M on the same evaluation, so the same task costs more.

How many parameters does Grok 4.6 have?

Undisclosed. Musk described a 2-trillion-parameter model on 18 July 2026 and confirmed the name, but SpaceXAI's release materials give no parameter count, no architecture details, and no statement on how many parameters activate per token.

What is Grok 4.6's context window and knowledge cutoff?

500K tokens of context, with text and image input and text-only output. The documentation gives a knowledge cutoff of 1 February 2026, and the model has no access to events after that without server-side Web Search or X Search enabled.

Is Grok made by xAI or SpaceXAI?

Both names refer to the same company. The developer documentation is titled "SpaceXAI Docs" and the announcement footer still reads X.AI LLC, after the 2026 move under SpaceX. The x.ai domains and the grok-* model IDs were not renamed.

When is Grok 4.7 coming out?

No date has been announced. Musk's 24 July post put Grok 4.7 about four weeks out, pointing at around 21 August 2026, and Grok 4.6 arrived five days past its own implied date.

Related reading

Pricing, model IDs and benchmark figures verified against SpaceXAI documentation, the Grok 4.6 announcement and Artificial Analysis on 13 August 2026.