Grok Imagine 2 is xAI's new image model, the Quality Mode of grok.com/imagine, GA on August 7, 2026. It is not a follow-up to the Grok Imagine 1.5 video model. It ranks #2 on the LM Arena image leaderboard in both text-to-image and image editing, behind GPT-Image-2. One catch worth knowing up front: its API isn't open yet, so generating or editing images programmatically today means using Grok's consumer app or switching to an API-accessible rival.
Grok Imagine 2 is an image model, not the video sequel
The single most useful thing to know about "Grok Imagine 2" is what it is: the image model Imagine Image 2.0 (model id grok-imagine-image-quality), which shipped as the Quality Mode of Grok's image generator. It is not the video model grok-imagine-video-1.5 that our existing Grok Imagine 1.5 review covers.
The naming collision is real and it trips people up. A quick disambiguation checklist: if a page lists native 4K resolution, 30-second clips, the Aurora video engine, or lip-synced spoken audio, those are video-model specs being attached to an image model. Image 2.0 produces still images and edits still images, with no duration, audio, or video timeline. Outputs land on grok.com/imagine and the iOS and Android apps.
xAI's announcement page is the cleanest source on the launch:
What Image 2.0 actually does
Imagine Image 2.0 is a generation-plus-editing image model built around layout fidelity and editing control. Its headline capability is producing images with legible text and dense, multi-element compositions (posters, product pages, infographics) without the letter-garbling that older image models were known for.
The capability set, per xAI's launch page (verified August 7, 2026):
| Capability | What it does |
|---|---|
| Magic-wand regional edit | Point at an area and change only that region (inpainting-style). |
| Segmentation | Make precise selections before editing, instead of painting a mask by hand. |
| Background removal | Remove backgrounds and export with transparency. |
| Multi-image reference edit | Feed up to 5 reference images in a single edit to keep a subject or style consistent. |
| Smart Resize | Generative outpainting across 9 aspect ratios (1:2, 9:16, 2:3, 3:4, 1:1, 4:3, 3:2, 16:9, 2:1); the model extends the canvas. |
| Text & layout fidelity | Small text and multi-element layouts stay legible and aligned. |
| 15 templates | Photo Edit, Product Color Change, Editorial Product Poster, Reimagine, Photo Collage, Mascot Maker, BG Removal & Change, E-Commerce Photos, UGC Photos, Professional Headshot, Icon Maker, Character Sprite, Props & UI Kit, Emoji Creator, Merch Maker. |
| Build a world for video | Generate characters, scenes, and props separately with a consistent style, a pre-production step for video. |
The throughline is editing control, the feature competitors benchmark Image 2.0 against, and the one we test below against its API-accessible rivals.
How good is it? The Arena ranking, read honestly
xAI says Imagine Image 2.0 is the world's #2 model in both text-to-image and image editing on LM Arena, listed under the name SpaceXAI (xAI's listing on the board). The #1 in both categories is GPT-Image-2.
Read this chart with the right expectations. The Elo numbers come from xAI's own reporting of public Arena data, not an independent test we ran, and the gap to GPT-Image-2 is small in editing (1,439 vs 1,463) and larger in text-to-image (1,320 vs 1,380). "#2" is a strong result, but it doesn't put Image 2.0 on par with the top model, and because the API isn't publicly open we couldn't run a controlled API head-to-head. To check the ranking yourself, the live leaderboard is at lmarena.ai; the next tier down includes Reve 2.1, Meta Muse-Image, Qwen-Image-3.0-Pro, Gemini Image, and Seedream.
Can you call the Grok Imagine 2 API today? No.
The API is not open. xAI's launch page states plainly, "API access is coming soon," with no date (checked August 10, 2026).
There's an apparent contradiction worth resolving. docs.x.ai already lists the model id grok-imagine-image-quality and the endpoints POST /v1/images/generations and POST /v1/images/edits (up to 3 reference images). Having a model id in the docs is not the same as having live, billable access. There is no published per-image price, and with access still "coming soon," grok-imagine-image-quality isn't callable yet. "Coming soon" wins over the docs listing.
That settles the relay question: AIReiter lists Grok only as the video model grok_imagine_1_5, so there's no AIReiter route to Grok images either. The only way to use Imagine Image 2.0 today is the Grok consumer app (grok.com, iOS, Android).
What you can use via API right now — tested
Since Grok's image API is closed, the practical question is what you can call today that sits on the same Arena leaderboard. Three of the top rivals are API-accessible: GPT-Image-2 (the #1), Nano Banana Pro (built on Gemini image tech), and Seedream V5 Pro (ByteDance). I ran both of Image 2.0's signature jobs, text-accurate generation and controlled editing, through them to see whether the capabilities that make Image 2.0 notable are exclusive to it.
Generation: same-prompt duel
I gave all three models the same prompt: a retro travel poster demanding a large title (NEON KYOTO), a sub-line (2049 EXPRESS), and a price line (TICKETS FROM 29000 YEN), in a muted teal-and-magenta palette. Same 1:1 ratio, default settings, one generation each. The point was to test text rendering and layout, Image 2.0's advertised strengths.
The result: all three models rendered all three text strings correctly (title, sub-line, and price line, with no spelling errors or hallucinated characters). At this sample size (n=1, one prompt, not a benchmark), you cannot rank the three by text ability, only confirm that all three rendered the requested text correctly on this poster prompt. That means the text-rendering capability Image 2.0 markets is reachable through API-accessible rivals for this kind of job, not that it is no longer distinctive across the board. They diverge on cost: GPT-Image-2 charged 1 credit, Nano Banana Pro 6, and Seedream V5 Pro 7 on the same job. If you want to try these directly, GPT-Image-2, Nano Banana Pro, and Seedream V5 Pro each have an API page.
Editing: controlled before/after with rivals
Editing is the other half of Image 2.0's pitch: keep the structure, change one thing. I fed the poster above into Nano Banana Pro as a reference image and asked it to keep the layout, all three text strings, Mount Fuji, and the train, while swapping the neon night city for a snowy Kyoto dawn with cherry blossoms in soft pink and gold.
The result: the three text strings all survived correctly (NEON KYOTO, 2049 EXPRESS, TICKETS FROM 29000 YEN), and Mount Fuji and the train are still there. The scene shifted from a dark neon night to a bright dawn. One honest caveat: the edit came back at a different aspect ratio (1376×768 vs the 1024×1024 original), so it reads more like a guided re-render than a strict, canvas-preserving inpaint. This was a single test of one model, so it shows the "preserve text and subject, change the scene" style of edit is achievable through an API-accessible rival today, not that every rival matches Grok's more surgical magic-wand inpaint (whose API is closed and we could not self-test; treat any official Grok editing demo as a vendor showcase, not something we verified).
Pricing: what an image costs through each route
Per-generation prices from each model's AIReiter page (linked in the table), verified August 10, 2026. Grok Image 2.0 has no API price because it has no API; it is available only in the Grok consumer app, with no published per-image figure.
| Model | On AIReiter | ~1K-image price (per generation) | Arena rank | Notes |
|---|---|---|---|---|
| GPT-Image-2 | Yes | $0.02 | #1 (T2I & Edit) | Cheapest of the three and the top-ranked. |
| Nano Banana Pro | Yes | $0.05 | Top tier | Gemini-based; strong text + editing. |
| Seedream V5 Pro | Yes | $0.07 (7 credits) | Top tier | ByteDance; no 4K tier; reference images $0.005 each. |
| Grok Image 2.0 | No (API pending) | n/a | #2 (T2I & Edit) | Consumer-only; available in the Grok app. |
Which way to go right now
Three routes, depending on what you need:
- You specifically want Grok Image 2.0. Use the Grok app (web/iOS/Android). There is no API path to it, and no date on when one opens.
- You want to generate or edit images through an API today. Use GPT-Image-2, Arena #1 at $0.02 per generation at 1K, or Nano Banana Pro / Seedream V5 Pro to compare. All three passed our one-generation text sample, and Nano Banana Pro held up on the editing test.
- You actually wanted video. That's a different model,
grok-imagine-video-1.5, covered in our Grok Imagine 1.5 review and on the Grok video page.
FAQ
Is Grok Imagine 2 an image or video model?
An image model. "Grok Imagine 2" refers to Imagine Image 2.0 (model id grok-imagine-image-quality), the Quality Mode of grok.com/imagine. The video model is the separate grok-imagine-video-1.5; if a description mentions 4K, 30-second clips, or audio, that's the video model, not this one.
Can I call the Grok Imagine 2 API?
Not yet. xAI says API access is "coming soon" (verified August 7, 2026). docs.x.ai lists the model id and the /v1/images/generations and /v1/images/edits endpoints, but there is no live, billable access and no per-image price published as of August 10, 2026.
How much does Grok Imagine 2 cost?
Through the API, there is no price because there is no API. In the Grok consumer app, xAI has not published a standalone per-image figure.
What's the best Grok Imagine 2 alternative via API?
GPT-Image-2. It ranks #1 on LM Arena (ahead of Grok Image 2.0) in both text-to-image and editing, and at roughly $0.02 per generation at 1K it is also the cheapest of the API-accessible top models we tested.