AIREITER

Claude Opus 5 vs GPT-5.6: Sol, Terra, Luna Tested and Priced

Last Updated: 2026-07-29 11:16:04

On July 29, 2026 we gave Claude Opus 5, GPT-5.6 Sol, Terra and Luna the same four tasks. Sixteen answers came back, and all sixteen were correct. On these four small checks, accuracy gives you nothing to choose between.

The models still separated sharply on two things you pay for: Opus 5 billed 1,642 output tokens where Sol billed 480, and above 272,000 input tokens GPT-5.6 requests jump to a higher rate card while Claude Opus 5 bills its full 1M window flat. At a 200K-input, 20K-output request shape the two flagships land within 7% on list price; the gap widens with output volume and past the 272K line, so the decision turns on request shape and behavior.

We ran all four models on the same four tasks

Each model received an identical user prompt for four checks with mechanically verifiable answers: fix a buggy Python median function, count letter occurrences in a fixed string, compute the 7.5-degree angle a clock shows at 3:15, and refactor a deliberately messy JavaScript function without changing behavior. We verified the code answers by executing them against edge cases, not by eyeballing.

Every model passed every check. The Python fixes all handled even-length lists correctly, all four said 9 and 7.5 where 9 and 7.5 were the answers, and all four refactors survived a behavioral diff that included falsy fields and null category filters.

Grouped bar chart of visible output tokens per task: Claude Opus 5 used the most on every task, peaking at 1,308 tokens on the refactor, versus 293 for GPT-5.6 Sol, 150 for Terra and 310 for Luna

Three caveats before you quote these numbers. This is one run per model and task, a smoke test rather than a benchmark. The calls went through a gateway that injects its own system prompt, so input tokens jumped between 35 and 4,415 across requests and we discarded them entirely. And both vendors bill hidden reasoning tokens as output, so "output tokens" here means what you are charged for, not what you read.

One qualitative note survived the tie: Opus 5 produced the strongest refactor of the four. It collapsed the original's two loops into one pass, extracted the 1.2 multiplier into a named constant, and left a comment explaining that it kept loose equality on purpose to preserve behavior. Sol, Terra and Luna all returned correct but more literal two-loop translations.

The difference is tokens and time, not correctness

Correctness was identical across 16 runs, so the working comparison is what each answer cost. Opus 5 finished the four tasks in 32.4 seconds total and billed 1,642 completion tokens. Sol took 19.7 seconds and 480 tokens, Terra 11.7 seconds and 308, Luna 15.3 seconds and 537.

Single run, July 29, 2026CorrectOutput tokensTotal timeOutput-only cost
Claude Opus 54/41,64232.4s$0.0411
GPT-5.6 Sol4/448019.7s$0.0144
GPT-5.6 Terra4/430811.7s$0.0046
GPT-5.6 Luna4/453715.3s$0.0032

The gap is reasoning volume, not answer length. Opus 5's refactor billed 1,308 tokens for a final answer of 492 bytes, and it spent 64 tokens to answer "7.5" where Sol spent 7. Anthropic's docs note that Opus 5 defaults to high effort on the API, so part of this is a default setting rather than a fixed trait; we measured default behavior, not a floor.

Community data matches this shape. A chart post on r/Anthropic, the first Google result for this query when we checked on July 29, shows the two nearly tied on coding scores while Opus takes longer and consumes significantly more tokens. One r/ClaudeCode user put the same finding less politely:

GPT 5.6 Sol just gets on with what you asked it to do. Opus 5 writes "War and Peace" in 12 volumes and then gets on with what you asked it to do.

Effort settings can flip the economics, though. A Hacker News commenter reading Artificial Analysis data noted that Opus 5 on high effort landed at roughly $1.06 per intelligence-index run against $1.04 for Sol on max, near parity once both models reason hard.

What "GPT-5.6" means: Sol, Terra and Luna

GPT-5.6 is not one model. OpenAI ships three SKUs, and the bare string gpt-5.6 is documented as an alias for gpt-5.6-sol, the most expensive of the three. A config that names gpt-5.6 bills at the family's top rate, which is worth checking before any cost comparison.

OpenAI's model reference describes Sol as its "frontier model for complex professional work," Terra as the tier that "balances intelligence and cost," and Luna as "optimized for cost-sensitive workloads." The three share a 1.05M-token context window, 128K maximum output, a February 16, 2026 knowledge cutoff, and six selectable reasoning efforts from none to max.

Spec, verified July 29, 2026Claude Opus 5GPT-5.6 (all three)
Context window1M tokens1.05M tokens
Max output128K (300K via Batch API beta)128K
Reliable knowledge cutoffMay 2026February 16, 2026
Reasoning controlAdaptive thinking, effort defaults to highnone / low / medium / high / xhigh / max
API IDclaude-opus-5gpt-5.6-sol / -terra / -luna

Two of those rows do real work in a decision. Opus 5's May 2026 knowledge cutoff is three months fresher than the whole GPT-5.6 family, which matters for anything referencing spring 2026 releases. And the reasoning controls differ in kind, not only in count: GPT-5.6 lets you switch reasoning off entirely with none, while Opus 5 manages its own thinking budget adaptively and exposes effort levels on top.

Pricing: a flat 1M window vs the 272K step

Both vendors publish clean per-token rates, and below 272K input tokens they are close: $5 per million input on both flagships, $25 per million output for Opus 5 against $30 for Sol, identical $0.50 cached-input rates, and a 50% batch discount on both sides. Priced at those rates, a 200K-input, 20K-output request costs $1.50 on Opus 5 and $1.60 on Sol.

Per 1M tokens, list price July 29, 2026InputCached inputOutputBatch
Claude Opus 5$5.00$0.50$25.00$2.50 / $12.50
GPT-5.6 Sol$5.00$0.50$30.00$2.50 / $15.00
GPT-5.6 Terra$2.50$0.25$15.00$1.25 / $7.50
GPT-5.6 Luna$1.00$0.10$6.00$0.50 / $3.00
OpenAI's API pricing page showing separate short-context and long-context rate columns for gpt-5.6-sol, terra and luna

The structural difference sits at 272K input tokens. OpenAI's pricing page splits every GPT-5.6 rate into short-context and long-context columns: past 272K input, input bills at 2x and output at 1.5x, which takes Sol to $10 and $45. Anthropic's pricing docs state the opposite design in one sentence: "A 900k-token request is billed at the same per-token rate as a 9k-token request."

Anthropic's pricing documentation listing Claude Opus 5 at $5 input and $25 output per million tokens

In dollars, with 20K output at each point: a 300K-input request costs $2.00 on Opus 5 and $3.90 on Sol. At 500K it is $3.00 against $5.90, and at 900K, $5.00 against $9.90. Above the step, Opus 5 runs at roughly half of Sol's price for the same shape of request.

Line chart of request cost by input size: Sol's curve breaks upward at 272K while Opus 5, Terra and Luna rise linearly; Terra tracks slightly below Opus 5 above the step

The chart holds one surprise that cuts against a clean Anthropic win. Above 272K, Terra's doubled input rate is $5, the same as Opus 5's, and its long-context output rate of $22.50 undercuts Opus 5's $25. At 20K output that makes Terra a constant five cents cheaper per request at 300K, 500K, 700K and 900K alike. The five cents scales with output volume, not input size: it is $2.50 per million output tokens of difference.

Three footnotes for budget planning. Anthropic sells a fast mode for Opus 5 in research preview at $10 input and $50 output; OpenAI sells a priority tier at 2x standard rates, short context only. US-region data residency adds about 10% at both vendors. And Claude's current tokenizer produces roughly 30% more tokens than pre-4.7 models for the same text by Anthropic's own note, so the honest final check is running one representative prompt through both tokenizers before you multiply rates.

Which model to call, by workload

The correctness tie means routing comes down to request shape and working style. Developers in r/ClaudeCode run Sol at high effort as a reviewer and QA agent because it is thorough and strong with a browser; others in r/codex route most work to the cheaper GPT tiers at medium reasoning and report the best cost-performance there.

Request shapeCallWhy
Under 272K input, output-heavy agent loopsClaude Opus 5$25 vs $30 output; strongest refactor judgment in our run
Under 272K, latency-sensitive or high-volumeGPT-5.6 TerraFastest and cheapest passing model we tested
Above 272K inputClaude Opus 5 or TerraSol's long-context rates double; Terra stays five cents under Opus 5
Review, QA and browser-driven checkingGPT-5.6 SolSix effort levels to tune review depth; validate on your own QA set
Throwaway classification and extractionGPT-5.6 Luna$1/$6 passed all four checks

Not everyone agrees Sol deserves the QA seat: one r/codex thread calls Sol a regression from GPT-5.5 while crediting Opus 5 with completing more tasks per plan. Defaults are cheap to test: all four models sit in AIReiter's catalog (Opus 5, Sol), so you can replay your own tasks side by side.

The one number to carry out of this comparison: 272,000 input tokens. Below it, choose on behavior, latency and the $5 output-rate difference. Above it, the price argument belongs to Anthropic against Sol, with Terra five cents under Opus 5 and Luna far below both.

FAQ

Is gpt-5.6 the same model as GPT-5.6 Sol?

Yes. OpenAI's model reference lists the bare gpt-5.6 string as an alias that resolves to gpt-5.6-sol, so it bills at Sol's rates: $5 input and $30 output per million tokens below 272K input, $10 and $45 above.

Does Claude Opus 5 charge more for long-context requests?

No. Anthropic bills the full 1M window at standard rates on the regular API, stating that a 900K-token request costs the same per token as a 9K one. Only the optional fast mode carries a premium, at $10 input and $50 output.

Which is cheaper, Claude Opus 5 or GPT-5.6?

By tier. Below 272K input, Opus 5 undercuts Sol by $5 per million output tokens and the models are otherwise identically priced; Terra and Luna are far cheaper than both. Above 272K, Opus 5 costs roughly half of Sol per request, while Terra stays marginally cheaper than Opus 5.

Is Claude Opus 5 better than GPT-5.6 for coding?

Both fixed, counted and refactored correctly in our July 29 run, and the r/Anthropic chart post that led Google results for this query on July 29 shows them nearly tied on coding scores. Opus 5 showed the best refactoring judgment but billed 3.4x Sol's output tokens at default settings, so the coding choice is an economics and verbosity choice.

What are the knowledge cutoffs for Opus 5 and GPT-5.6?

Anthropic gives Claude Opus 5 a reliable knowledge cutoff of May 2026. All three GPT-5.6 models share a February 16, 2026 cutoff, about three months earlier.

Further reading