AIREITER

GPT-5.6 Luna vs GPT-5.4 Mini: Which Is Cheaper Now?

Last Updated: 2026-07-31 03:25:55

On July 30, 2026, OpenAI cut GPT-5.6 Luna's API price by about 80%, taking it from the launch price of $1/$6 down to $0.20 input / $1.20 output per million tokens. At that price Luna is roughly 73% cheaper than GPT-5.4 mini ($0.75/$4.50), yet it matched mini answer-for-answer on both coding and reasoning prompts we ran. Mini still wins one place: it returned about twice as fast, so latency-bound workloads are the single scenario where the older small model still earns the call.

The July 30 price cut flipped this comparison

GPT-5.6 Luna reached general availability on July 9, 2026 at $1/$6 per million tokens, positioned as the cost-efficient tier inside the GPT-5.6 family but still well above OpenAI's cheap gpt-5.4-mini ($0.75/$4.50). Roughly three weeks later, on July 30, OpenAI dropped Luna about 80% to $0.20/$1.20 and trimmed GPT-5.6 Terra about 20% to $2/$12. Luna now sits on the same shelf as gpt-5.4-nano ($0.20/$1.25): a nano-priced model with a current-generation brain.

If you have seen Luna quoted at $1/$6, that launch price stopped being live on July 30, 2026; the current Standard-tier price is $0.20/$1.20, verified July 31, 2026.

Cost reversal at a 10M-input plus 1M-output workload: Luna versus mini versus nano

Luna vs mini at list price

At current Standard-tier prices, Luna is cheaper than mini on both input and output, and the gap holds whether your prompts are short or long. Luna also carries a 1,050K context window against mini's 400K.

Model (Standard, per 1M tokens)InputCached inputOutputContext
gpt-5.6-luna (short context)$0.20$0.02$1.201,050K
gpt-5.6-luna (long context, >200K)$0.40$0.04$1.801,050K
gpt-5.4-mini$0.75$0.075$4.50400K

Source: developers.openai.com/api/docs/pricing, verified 2026-07-31. For reference, gpt-5.4-nano lists at $0.20/$1.25, the price band Luna now shares.

Run the math at a moderate workload of 10M input plus 1M output and Luna costs $3.20 against mini's $12.00 (nano is $3.25). Drop to 1M plus 1M and it is $1.40 against $5.25. Luna does charge double above 200K tokens of input (its long-context rate is $0.40/$1.80), but even that band undercuts mini's flat $0.75/$4.50. Mini has no long-context tier at all; its 400K window is a single short-context price band. For a deeper breakdown of the whole family's pricing, see the GPT-5.6 pricing guide.

OpenAI developer pricing page showing the gpt-5.6-luna row at $0.20 input and $1.20 output, verified 2026-07-31

Same prompt, both models: quality, tokens, and speed

We sent identical prompts to both models through the same API endpoint: a coding task (return the length of the longest valid parentheses substring) and a reasoning task (the snail-in-a-well problem, run three times). Both models solved both correctly, producing the same stack-based algorithm for the parentheses problem and "28 days" on the snail on each of the three runs. They diverged on speed and verbosity, not correctness.

Quality and correctness

On the coding prompt, Luna and mini produced functionally identical O(n) stack solutions with a sentinel base index; neither made an error. On the reasoning prompt, both reached the correct 28 days and both correctly noted that the snail escapes on day 28 before it can slip back. On these two basic tasks the two are effectively tied; we did not stress-test harder reasoning.

Speed and latency

Mini was about twice as fast end-to-end. Across three reasoning runs, mini averaged ~2.8s against Luna's ~5.5s; on the single coding run, mini took 4.0s against Luna's 4.6s. Luna's extra time tracks its more verbose output, 168 tokens against mini's 76 on the reasoning task. It favors completeness over snap replies.

"It's basically as fast as 5.4 mini if you use medium thinking and reasons really well, like above 5.4 level." — r/codex, "GPT-5.6 Luna is really underrated" thread

At default settings, mini is the faster return.

Per-task cost

Luna's greater verbosity is more than offset by its lower unit price. On the reasoning prompt, Luna emitted an average of 168 output tokens against mini's 76, yet at current prices that response cost roughly $0.001 on Luna versus $0.004 on mini. These are sub-cent, single-digit-sample figures (run 2026-07-31 via the AIReiter API on gpt-5.6-luna and gpt-5.4-mini, default effort, US endpoint; times are full-response wall-clock, token counts are API-reported completion tokens), so treat them as directional rather than as a benchmark. The direction is unambiguous: per correct answer, Luna is cheaper.

Why the "which is smarter" benchmark fight is mostly a settings mismatch

On raw scorecard numbers Luna beats mini across coding and reasoning: SWE-Bench Pro 62.7 vs 54.4, GPQA Diamond 92.3 vs 88.0, Toolathlon 53.4 vs 42.9 (scores from OpenAI's model scorecards). The one widely-cited result that flips the verdict comes from Artificial Analysis's Intelligence Index, where GPT-5.4 scores 51 to Luna's 46. But that index compares Luna run at high effort against GPT-5.4 run at xhigh effort. That is a different reasoning budget on each side, not a level playing field.

BenchmarkGPT-5.6 LunaGPT-5.4 miniNote
SWE-Bench Pro62.754.4Luna higher
GPQA Diamond92.388.0Luna higher
MMMU Pro78.476.6Close
Toolathlon53.442.9Luna higher
Terminal-Bench84.7 (v2.1)60.0 (v2.0)Different versions
OSWorld45.6 (v2.0)72.1 (Verified)Different protocols

Scores compiled from OpenAI's published model scorecards and cross-checked against docsbot.ai.

Two caveats matter when you read these. First, effort tier: mini exposes a fine-grained None / Low / Medium / High / XHigh dial, while Luna defaults to a high reasoning tier, so turning mini up to xhigh and comparing it to Luna at its default inflates mini's headline. Second, the rows are not strictly like-for-like. docsbot.ai flags that the Terminal-Bench and OSWorld figures mix benchmark versions and protocols, which is how mini's 72.1 OSWorld can look higher than Luna's 45.6 while testing something subtly different.

Beyond the scores, Luna carries a February 2026 knowledge cutoff against mini's August 2025, and a 1.05M context window against mini's 400K.

The "Luna costs 96% more" thread is now obsolete

A widely-circulated OpenAI community test found Luna cost 96% more than mini per six-turn conversation. That number is genuine but pre-price-cut: it was measured on July 11 against Luna's $1/$6 launch price, and most of the gap came from a cache-invalidation bug in the test app rather than from the model itself. Once the author stopped mutating the top-level instructions each turn, Luna's premium fell to 16.7%, and Luna was 9.9% faster with a higher blind-rated quality score (4.80 vs 4.28).

At the current $0.20/$1.20 price, Luna is cheaper than mini before caching even enters the picture, so the cache argument no longer moves the Luna-vs-mini decision. The mechanism is still worth knowing for cost tuning: OpenAI bills cache writes at 1.25× the input rate ($0.25/M for Luna now) and cache reads at a 90% discount, with a 30-minute minimum cache life. If your app rewrites the system prompt each turn, the cache prefix invalidates and you pay the write rate instead of the read rate. That is the trap that produced the 96% figure.

Which one to call

For most workloads the answer is now Luna: cheaper at both context tiers, matched or better on quality, with a newer knowledge cutoff and a 2.6× larger window. Reach for mini when latency is the hard constraint, or when you need effort dialing that Luna does not expose.

If you want…PickWhy
Lowest cost per correct answerLuna$0.20/$1.20, ~73% below mini
Lowest latencymini~2× faster end-to-end (~2.8s vs ~5.5s on reasoning)
Fine-grained None/XHigh effort controlminiLuna is high-tier only
Prompts over 400K tokensLuna1.05M window; mini caps at 400K
The most concise outputminiEmitted ~half the tokens on equal tasks

A practical routing split: mini as the default where its latency advantage matters, Luna as the escalation and long-context model where quality and window size matter. Both are callable through a unified API as gpt-5.6-luna and gpt-5.4-mini, for example on the AIReiter GPT-5.6 Luna endpoint. One routing gotcha: the flagship Sol prices at $5/$30, about 25× Luna, so if your endpoint aliases a bare gpt-5.6 id to Sol, pin the full gpt-5.6-luna id.

FAQ

Is GPT-5.6 Luna cheaper than GPT-5.4 mini?

Yes, about 73% cheaper at current pricing. After the July 30, 2026 cut, Luna is $0.20 input / $1.20 output per million tokens against mini's $0.75 / $4.50, on both input and output.

Is GPT-5.6 Luna faster or slower than 5.4 mini?

Slower, by roughly 2× in our same-prompt test (mini averaged ~2.8s against Luna's ~5.5s on reasoning). Luna trades speed for more complete reasoning; turning its effort down narrows the gap.

Does GPT-5.6 Luna replace the mini tier?

No. Luna is the cost-efficient tier inside the GPT-5.6 generation; gpt-5.4-mini is a separate, older small model that is still sold and supported.

GPT-5.4 nano vs Luna: which is cheaper?

Effectively tied. Luna is $0.20/$1.20 against nano's $0.20/$1.25. Luna is fractionally cheaper on output and one generation newer, with a later knowledge cutoff.

When were Luna and mini released?

Luna launched July 9, 2026 and was cut to its current price on July 30, 2026. GPT-5.4 mini launched March 17, 2026.

Sources