After three weeks of "coming soon," Qwen3.8-2.4T-A95B weights went live on ModelScope on August 12, 2026. The headline number is 2.4 trillion total parameters - but only 95 billion activate per token. The catch: even at 4-bit quantization, the model occupies roughly 1.2 TB of VRAM, and almost nobody outside a datacenter can run it locally. Here is what the release contains, what it costs to use, and when the API is the only sensible path.
Qwen3.8-2.4T-A95B Weights Are Now Downloadable
The weights are live. Alibaba's Qwen team published the checkpoint for Qwen3.8-2.4T-A95B through its ModelScope community account late on August 12, making this the first Qwen-Max-tier model to receive an open-weight release. All three previous Qwen Max models - Qwen3.5-Max, Qwen3.6-Max, Qwen3.7-Max - stayed API-only.
One naming distinction matters immediately. There are two things called "Qwen3.8":
| Name | What it is | Where to get it |
|---|---|---|
| Qwen3.8-2.4T-A95B | Open-weight base model | ModelScope (downloadable checkpoint) |
| Qwen3.8-Max | Cloud-hosted version with additional production capabilities | Alibaba Token Plan, Qoder, QoderWork |
Qwen3.8-Max is built on the open-weight A95B model but adds post-training enhancements. Alibaba describes a three-stage system spanning real-environment scaling, unified reward signals, and online data balancing. The open weights give you the base model, not the full Max production stack.
Guides that say the weights are "promised but unavailable" are now outdated - the release landed within Alibaba's projected "week of August 10" window.
Confirmed Specs: 95B Active, 256K Native Context
For weeks after the July 19 preview launch, the active-parameter count was the single most-requested missing spec. Multiple third-party guides explicitly listed it as "not published." Now it is confirmed: 95 billion active parameters per token out of 2.4 trillion total.
| Specification | Confirmed Value |
|---|---|
| Total parameters | 2.4T |
| Active parameters per token | 95B |
| Architecture | MoE (continuing Qwen3.5's hybrid design) |
| Experts per MoE layer | 512 |
| Routed experts per token | 10 |
| Shared experts per token | 1 |
| Native context window | 262,144 tokens (256K) |
| Expandable context | 1,010,000 tokens (1.01M) |
| Training method | Multi-Token Prediction (MTP) |
| First open-weight Qwen-Max release | Yes (August 12, 2026) |
Two details deserve attention. First, the native context window is 256K, not 1 million tokens as widely assumed. The 1.01M figure refers to an expandable mode whose quality tradeoffs Alibaba has not yet documented. Second, the MoE routing is extremely sparse - only 11 of 512 experts fire per token (10 routed plus 1 shared). That ~4% parameter activation ratio (95B of 2.4T) is what makes $2/M input economically viable despite the 2.4T total footprint.
MTP (Multi-Token Prediction) training means compatible inference engines can predict multiple tokens per forward pass, potentially improving throughput. Alibaba has not specified which engines support this.
Self-Hosting Reality: What 2.4T Parameters Actually Costs
The specs sound impressive until you try to load them. Here is the arithmetic.
At 4-bit quantization, 2.4 trillion parameters consume approximately 1.2 TB of storage - before accounting for KV cache, activation memory, or runtime overhead. An NVIDIA H200 GPU has 141 GB of memory. Eight H200s give you 1,128 GB total, which is not enough to hold the weights and serve any meaningful context window simultaneously.
| Deployment Scenario | VRAM Needed (est.) | Hardware | Practical? |
|---|---|---|---|
| Full precision (16-bit) | ~4.8 TB | 34+ H200 GPUs | Datacenter only |
| 4-bit quantized | ~1.2 TB | 9+ H200 GPUs | Datacenter only |
| 1.5-bit quantized | ~450 GB | 4+ H200 GPUs | Tight; quality loss uncertain |
| Qwen3.8-27B (alternative) | ~27 GB at 4-bit | 1 consumer GPU | Yes |
The 95B active-parameter count helps inference throughput per token - the model only routes through a fraction of its experts on each forward pass. But active parameters do not reduce the memory needed to store all 512 expert networks. As GEO Toolbox put it: "Sparse activation makes the model cheap to run at scale, not cheap to own."
For local or single-server deployment, the Qwen3.8-27B variant is the realistic option. It shares the Qwen3.8 training lineage but fits on a single consumer GPU at 4-bit. The 2.4T model is a research and datacenter play.
When the API Is the Smarter Path
Since self-hosting a 2.4T MoE model is out of reach for nearly every team, the API is where practical access happens. Alibaba's Model Studio lists Qwen3.8 Max at $2 per million input tokens and $6 per million output tokens - cheaper than both its predecessor and its closest Chinese rival.
| Model | Input ($/M) | Output ($/M) | Context |
|---|---|---|---|
| Qwen3.8 Max | $2.00 | $6.00 | 256K native (1.01M expandable) |
| Qwen3.7 Max | $2.50 | $7.50 | 1M |
| Kimi K3 | $3.00 | $15.00 | 1M |
Beyond the per-token API, Alibaba offers three subscription tiers:
- Lite: $6/month (~10,000 credits)
- Standard: $18/month (~40,000 credits)
- Pro: $68/month (~160,000 credits)
The Token Plan endpoint supports both OpenAI-compatible and Anthropic-compatible API formats, so tools like Claude Code, Cursor, and OpenCode work without client-side changes. Qoder, Alibaba's coding product, offers promotional credit rates as low as 0.01x standard during off-peak hours (14:00-00:00 UTC) - but these are launch discounts, not durable pricing.
The subscription model has a sharp edge. Qwen3.8 Max's reasoning mode is verbose, and community reports of quota burn are consistent:
"Qwen tends to overthink A LOT." - u/darko, r/Qwen_AI
In the same Reddit thread, one Lite subscriber reported exhausting their weekly allowance in under two days on roughly 50 million tokens. A Pro-plan user described burning 500 million tokens in a single day across seven parallel agents with a 94% cache-hit rate. The reasoning_effort parameter (settable to low, medium, or xhigh) directly controls how much reasoning output the model generates - and therefore how fast credits disappear. Setting it to low reduced thinking time to roughly 10 seconds per request in one user's test.
For a detailed cost breakdown, see our Qwen3.8 Max API pricing guide.
Benchmarks: Where Qwen3.8 Max Wins and Falls Short
Alibaba published benchmark numbers comparing Qwen3.8 Max against Fable 5 and GPT-5.6 Sol. The picture is mixed - and the gap between vendor-reported and independent scores is large enough to matter.
| Benchmark | Qwen3.8 Max (Alibaba) | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| Terminal-Bench 2.1 | 86.6 | 84.6 | Higher than Qwen (exact score not published) |
| SWE-bench Pro | 67.7 | 80.0 | 64.6 |
| GPQA Diamond | 92.6 | 92.6 | Higher than Qwen |
| Humanity's Last Exam | 43.6 | 53.3 | Not published |
The independent picture tells a different story. Artificial Analysis measured Qwen3.8 Max at 81.3% on Terminal-Bench 2.1 - a 5.3-point drop from Alibaba's 86.6%. At that independent score, Qwen ranks roughly tenth overall, behind Fable 5 (84.6%) and Kimi K3 (85.0%).
Where Qwen3.8 Max does hold an edge: SWE-bench Pro at 67.7% beats GPT-5.6 Sol's 64.6%, and multimodal scores are strong (MathVision 95.2 vs Fable 5's 92.7, LogicVista 91.9 vs 85.7). The model also improved sharply from Qwen3.7 Max on DeepSWE 1.1 (21.6 -> 56.6), suggesting real gains in software-engineering agent tasks.
Community sentiment reflects the split. The same Reddit thread that complained about overthinking also had this:
"The output was god tier though." - u/TangerineLogical9779, r/Qwen_AI
Speed is reportedly unstable, ranging from 8 tokens/sec to 60 tokens/sec depending on load and time of day. For a head-to-head breakdown against Kimi K3, see our Qwen3.8 Max vs Kimi K3 comparison.
FAQ
Are Qwen3.8-Max weights available to download?
Yes. Alibaba released the Qwen3.8-2.4T-A95B checkpoint on ModelScope on August 12, 2026. This is the first Qwen-Max-tier model with open weights. The cloud-hosted Qwen3.8-Max includes additional production capabilities on top of this base model.
What license do the open weights use?
The exact license has not been confirmed at the time of writing. Historical precedent suggests Apache 2.0 - Qwen3.6 shipped under that license - but Qwen3.7 Max remained closed. Read the actual license file on the ModelScope repository before building anything commercial.
How is Qwen3.8-2.4T-A95B different from Qwen3.8-Max?
Qwen3.8-2.4T-A95B is the open-weight base model: 2.4T parameters, 95B active, MoE architecture, 256K native context. Qwen3.8-Max is Alibaba's cloud version built on this model, augmented with production-environment post-training including unified reward systems and online data balancing. The open weights give you the base; the API gives you the enhanced version.
Can I run Qwen3.8-Max locally?
Not practically. At 4-bit quantization, the model requires approximately 1.2 TB of VRAM - nine or more H200 GPUs for weights alone, with no room for context cache. Even 1.5-bit quantization leaves hundreds of gigabytes. The Qwen3.8-27B variant is the realistic local alternative, fitting on a single consumer GPU.
How much does the Qwen3.8 Max API cost?
The Model Studio API charges $2 per million input tokens and $6 per million output tokens. Subscription tiers start at $6/month (Lite) and go up to $68/month (Pro). Qoder offers promotional credit discounts up to 98% off during off-peak hours, but these are launch pricing and may change.