AIREITER

OpenAI Ultrafast Mode: 14x GPT-5.6 Sol Speed, No Price Yet

Last Updated: 2026-08-17 00:33:05

Three OpenAI offerings now have "fast" or "ultra" in the name, and only one of them runs at 750 tokens per second. OpenAI Ultrafast mode - announced August 13, 2026 - is a new API service tier that runs GPT-5.6 Sol on Cerebras wafer-scale chips at up to 14x Standard speed. The catch: access is invite-only, and OpenAI published no pricing for it, so the fastest tier you can actually buy today remains Fast mode at $10/$60 per million tokens.

The Three-Name Problem: ultra, Fast Mode, and Ultrafast

GPT-5.6 Sol buyers currently face three speed-flavored labels that mean different things. Here is the full disambiguation in one table.

NameWhat it isHow you get itSol price (per 1M tokens)Advertised speed
ultra reasoning settingA reasoning/working mode that spends subagents and extra compute on one taskCodex setting, not an API tier you purchase separatelyNo standalone rate; burns usage limits fastSlower by design
Fast modeAPI processing tier, renamed from Priority on July 30, 2026service_tier="fast" in any API request$10 in / $60 out (short context)Up to 2.5x Standard
Ultrafast modeCerebras-powered API tier for GPT-5.6 Sol, previewed August 13, 2026Limited preview, select customers onlyNot disclosedUp to 14x Standard

They move in opposite directions: ultra is the subagent working mode that trades speed for thoroughness, while Ultrafast mode is raw decode speed on the same model. Fast mode sits between them as a purchasable latency upgrade, documented separately on OpenAI's Fast mode page.

What the August 13 Preview Actually Ships

OpenAI Ultrafast mode announcement page with the GPT-5.6 Sol speed preview

The official announcement commits to five concrete facts. Ultrafast launches first in the OpenAI API, runs GPT-5.6 Sol only, delivers up to 14x Standard-processing speed and up to 750 output tokens per second, is powered by Cerebras, and ships as a limited preview for a select group of customers that includes Jane Street, Podium, Basis, and Rogo.

The announcement omits four things any buyer needs: a per-token price, a model ID, rate limits, and a latency SLA. No named benchmarks appear either.

OpenAI's evidence is a side-by-side demo of Ultrafast and Standard Sol building a 3D warehouse simulator from one prompt. One quote from Rogo's Alex Wang sums up the pitch: "Ultrafast makes complex financial research feel like a real-time interaction."

From 2.5x to 14x: What the Speed Gap Buys

Bar chart comparing advertised speed multiples of Standard, Fast mode, and Ultrafast for GPT-5.6 Sol

One measurement caveat first: 750 tokens per second is advertised output throughput, while Fast mode's number is an SLA floor - related but not identical metrics. Dividing 750 by 14 implies Standard Sol decodes at roughly 54 tokens per second, which puts a 10,000-token generation at approximately 186 seconds on Standard, up to 125 seconds at Fast mode's Enterprise SLA floor of 80 tokens per second, and about 13 seconds on Ultrafast.

WorkloadStandard (~54 tok/s, derived)Fast mode (SLA ≥80 tok/s)Ultrafast (750 tok/s advertised)
10,000-token generation~186 s≤125 s~13 s
5-turn agent loop, 1,000 tokens/turn~93 s≤63 s~7 s

Figures circulating on r/AIToolsPerformance attribute these comparisons to Artificial Analysis and Cerebras: Ultrafast at 11x the speed of Claude Fable 5 and 5x Opus 4.8 on Fast mode, and a full 2,500-question Humanity's Last Exam run in 11 hours 11 minutes versus 78 hours 27 minutes for Fable 5. All of these are vendor-reported; no independent replication exists yet.

The Price Nobody Outside the Preview Knows

OpenAI Fast mode pricing table page for API customers

Ultrafast pricing does not exist in public. The announcement names no rate, and community reaction zeroed in on that gap immediately - as one r/AIToolsPerformance poster put it: "What the blog doesn't say is price. Nobody outside the preview knows what it costs."

The known GPT-5.6 Sol price ladder, all per million tokens, is the anchor set available today:

TierInput (short ctx)Output (short ctx)Input (long ctx >272k)Output (long ctx)Cached input
Standard$5.00$30.00$10.00$45.00$0.50
Fast mode$10.00$60.00$20.00$90.00$1.00
Ultrafastundisclosedundisclosedundisclosedundisclosedundisclosed
Grouped bar chart of GPT-5.6 Sol Standard versus Fast mode prices per million tokens

Fast mode charges exactly 2x Standard for Sol; public information does not support estimating Ultrafast pricing from that multiple (rates can move both ways - OpenAI cut Terra/Luna prices on July 30). Fast mode's fine print matters for any latency budget: if you exceed 1 million tokens per minute and increase traffic by more than 50% in under 15 minutes, excess requests silently downgrade to Standard, return service_tier="Default", and bill at Standard rates.

Getting In: What Access Looks Like Today

Ultrafast access runs through invitation, not a public waitlist. OpenAI says the preview expands "as capacity grows" and offers an access-update signup on the announcement page, and July reporting counted only about 20 organizations inside early Sol access.

The infrastructure has precedent: OpenAI's Cerebras partnership dates to January 14, 2026 and reserves 750 MW of wafer-scale capacity, and GPT-5.3-Codex-Spark already exceeded 1,000 tokens per second on the same hardware.

Shipping Fast Output Before Your Invite Arrives

Fast mode is available to any API customer now, and enabling it is a one-parameter change:

  1. Add service_tier="fast" to a request on /v1/responses or /v1/chat/completions.
  2. Legacy service_tier="priority" values keep working - existing models still report priority in usage dashboards, while models after GPT-5.6 will show fast.
  3. To default an entire project, set Project settings -> Default Service Tier -> Fast.
Fast-mode modelPrice per 1M (in / out)Latency SLA
GPT-5.6 Sol$10 / $60>80 tok/s
GPT-5.6 Terra$4 / $24>70 tok/s
GPT-5.6 Luna$0.40 / $2.40>100 tok/s

Luna's SLA floor is the highest of the three. Cache reads cut 90% off input - $0.50/M on Standard Sol with a roughly 30-minute TTL - so a long-lived session beats fresh requests on any tier. To test these tiers against real traffic, the GPT-5.6 API endpoint exposes all three models through one interface.

FAQ

Is Ultrafast mode available in ChatGPT or Codex?

No. Ultrafast launches first in the OpenAI API, and the announcement gives no date for ChatGPT or Codex availability.

Does running GPT-5.6 Sol on Cerebras change output quality?

OpenAI positions Ultrafast as the same GPT-5.6 Sol served faster, not a different model. A Cerebras-reported 5.6x GDP-Val speedup with no quality drop is the only quality claim so far, and no independent benchmark has replicated it.

Is there a latency SLA for Ultrafast mode?

None is published. Fast mode carries a >80 tokens/sec p50 SLA and 99.9% uptime for Enterprise customers; Ultrafast has no stated SLA during the preview.

Which models support Ultrafast mode?

Only GPT-5.6 Sol. Terra and Luna have Fast-mode support but are not part of the Ultrafast preview.

How much will Ultrafast mode cost?

No official figure exists. Fast mode's published 2x premium over Standard Sol is the only reference point.