Qwen3.8-Max is now generally available at $2 per million input tokens and $6 per million output, and that output rate undercuts Claude Opus 4.8, Claude Fable 5 and GPT-5.6 Sol by four times or more. But reasoning_effort ships set to xhigh, so the sticker price and the invoice diverge from the first request.
What Qwen3.8-Max costs per token
Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model with 95 billion parameters active per token, and it takes over from Qwen3.7-Max as Alibaba's flagship. These are the rates published on Alibaba's model page, which carries its own "Updated: Aug 2, 2026" stamp:
| Item | Price per 1M tokens |
|---|---|
| Input | $2.00 |
| Output | $6.00 |
| Input (implicit cache) | $0.25 |
| Explicit cache creation | $2.50 |
| Explicit cache read | $0.17 |
The implicit cache tier needs no code. Repeated prefixes are billed at $0.25 automatically, an eighth of the standard input rate. Explicit caching is the tier that needs a decision: at $2.50 to write and $0.17 to read, it beats uncached input from the second read onward, but it does not overtake the $0.25 implicit tier until roughly the thirtieth read of the same prefix. For most applications that means leaving explicit caching alone.
One limit on all of this: the API is open, the weights are not. Alibaba's launch post says Qwen3.8-Max weights ship on Hugging Face and ModelScope next week, which is the first calendar commitment after weeks of an undated "coming soon", and as of August 3 the model page sidebar still reads Open Source: No.
Why the $6 output price understates your bill
Alibaba's launch documentation gives reasoning_effort three settings and names xhigh the default: the deepest of the three, above medium ("balances accuracy and speed") and low ("optimized for speed and cost"). preserve_thinking is also on by default for all scenarios. Thinking tokens bill as output tokens, which means both defaults route spend to the $6 rate before you have written a single prompt.
So the two levers on effective cost are set against you out of the box, and both are one request parameter away. Dropping to medium or low is the difference between paying $6 per million for deep deliberation you may not need and paying it only where the task earns it.
One early tester was blunt about what that deliberation buys, in an r/LocalLLaMA thread title:
Qwen 3.8 max first impressions - model is moderate and it is awful on thinking
Read that as a reason to test at each effort level on your own workload before standardising on the default, not as a benchmark result.
No long-context surcharge, and where that matters
The Qwen3.8-Max model page lists one input rate and one output rate, with no context-tier rows at all. Gemini 3.1 Pro and GPT-5.6 Sol both switch to a more expensive band above a context threshold:
| Model | Input (short ctx) | Output (short ctx) | Input (long ctx) | Output (long ctx) |
|---|---|---|---|---|
| Qwen3.8-Max | $2.00 | $6.00 | $2.00 (to 1M) | $6.00 (to 1M) |
| Gemini 3.1 Pro | $2.00 (≤200K) | $12.00 | $4.00 (>200K) | $18.00 |
| GPT-5.6 Sol | $5.00 | $30.00 | $10.00 | $45.00 |
Gemini 3.1 Pro matches Qwen3.8-Max exactly on input price up to 200K tokens, then doubles it. GPT-5.6 Sol doubles input and adds half again to output once a request crosses into its long-context band, on top of an already restructured price sheet.
So the gap widens with context length instead of holding steady. A 400K-token prompt costs $0.80 of input on Qwen3.8-Max, $1.60 on Gemini 3.1 Pro, and $4.00 on GPT-5.6 Sol. Document pipelines, long agent sessions, and whole-repository passes are where that compounds. On short chat turns it barely registers.
Cost against the models Alibaba benchmarked it against
Against the field, $6 of output is under a quarter of Claude Opus 4.8's $25 and a fifth of GPT-5.6 Sol's $30. Claude Fable 5, at $50, is more than eight times the rate. Whether that gap is worth taking depends on how the capability lines up, and Alibaba published a 28-row benchmark table to argue that it does:
It leads on PaperBench (93.0, against 90.5 for GPT-5.6 Sol and 88.8 for Fable 5), on JobBench (53.4 against Opus 4.8's 48.4), and edges both Claude models on Terminal Bench 2.1 at 86.6. It loses SWE-bench Pro badly: Fable 5's 80.0 is more than twelve points clear of its 67.7. Every one of these figures is vendor-run and published by the vendor on that same page.
That splits the decision cleanly. Output-heavy agentic work (long tool-calling loops, document generation, the research-reproduction shape PaperBench measures) is where a 4-to-8x output discount changes the unit economics of running the thing at all. Repository-scale software engineering of the kind SWE-bench Pro probes is where the discount buys you a real capability drop, and where paying Fable 5 rates still makes sense.
Testing that on your own workload is cheap to arrange: Alibaba's launch post documents OpenAI-compatible chat-completions and responses calls plus an Anthropic-compatible interface, which are the same two request shapes Claude Opus 4.8 and GPT-5.6 Sol accept.
What to change in your code today
The model id is qwen3.8-max. As of August 3 it is the only 3.8 entry in Alibaba's model catalogue — the preview identifier qwen3.8-max-preview is gone — so check your relay provider before assuming a preview discount survived GA.
Either interface is served from three regional base URLs: Beijing, Singapore, and US Virginia. The ceilings you will hit first:
| Limit | Value |
|---|---|
| Context window | 1M |
| Max input | 991.80K |
| Max input (thinking) | 983.61K |
| Max output | 131.07K |
| Requests per minute | 15K |
| Tokens per minute | 2M |
Note the roughly 8K-token gap between the two input ceilings: turning thinking on costs you input headroom as well as output tokens. And if next week's weights do land, the 95B active-parameter figure is where any self-hosting comparison against $2/$6 will start.
FAQ
Is the Qwen API free?
No. Qwen3.8-Max is billed per token at $2 input and $6 output per million. The Token Plan subscription sold on Alibaba's API platform is a separate credit-based product from per-token API billing.
What does Qwen Max cost?
For the current generation, $2 per million input tokens and $6 per million output, with implicit prefix caching at $0.25. Older qwen-max entries in Alibaba Cloud's documentation refer to previous generations and carry different rates.
Does Qwen3.8-Max charge extra for the 1M context window?
No. Input and output are priced the same at 5K tokens and at 900K, which is the clearest structural difference from GPT-5.6 Sol and Gemini 3.1 Pro.
Is Qwen3.8-Max cheaper than Qwen3.7-Max?
Not answerable from the official pages. The Qwen3.7-Max model page renders its per-token figures as -- behind a 50%-off marker rather than showing rates, so there is no published number to compare against.
Related reading
- Claude Fable 5 pricing: the $10/$50 tier this compares against
- Cutting Claude Fable 5 token costs: the same effort-and-caching levers on a different vendor