AIREITER

Qwen 3.8 Max Open Weights: Specs, API Pricing & Hardware

Last Updated: 2026-08-12 19:05:36

After three weeks of "coming soon," Qwen3.8-2.4T-A95B weights went live on ModelScope on August 12, 2026. The headline number is 2.4 trillion total parameters - but only 95 billion activate per token. The catch: even at 4-bit quantization, the model occupies roughly 1.2 TB of VRAM, and almost nobody outside a datacenter can run it locally. Here is what the release contains, what it costs to use, and when the API is the only sensible path.

Qwen3.8-2.4T-A95B Weights Are Now Downloadable

The weights are live. Alibaba's Qwen team published the checkpoint for Qwen3.8-2.4T-A95B through its ModelScope community account late on August 12, making this the first Qwen-Max-tier model to receive an open-weight release. All three previous Qwen Max models - Qwen3.5-Max, Qwen3.6-Max, Qwen3.7-Max - stayed API-only.

One naming distinction matters immediately. There are two things called "Qwen3.8":

NameWhat it isWhere to get it
Qwen3.8-2.4T-A95BOpen-weight base modelModelScope (downloadable checkpoint)
Qwen3.8-MaxCloud-hosted version with additional production capabilitiesAlibaba Token Plan, Qoder, QoderWork

Qwen3.8-Max is built on the open-weight A95B model but adds post-training enhancements. Alibaba describes a three-stage system spanning real-environment scaling, unified reward signals, and online data balancing. The open weights give you the base model, not the full Max production stack.

Guides that say the weights are "promised but unavailable" are now outdated - the release landed within Alibaba's projected "week of August 10" window.

Confirmed Specs: 95B Active, 256K Native Context

For weeks after the July 19 preview launch, the active-parameter count was the single most-requested missing spec. Multiple third-party guides explicitly listed it as "not published." Now it is confirmed: 95 billion active parameters per token out of 2.4 trillion total.

SpecificationConfirmed Value
Total parameters2.4T
Active parameters per token95B
ArchitectureMoE (continuing Qwen3.5's hybrid design)
Experts per MoE layer512
Routed experts per token10
Shared experts per token1
Native context window262,144 tokens (256K)
Expandable context1,010,000 tokens (1.01M)
Training methodMulti-Token Prediction (MTP)
First open-weight Qwen-Max releaseYes (August 12, 2026)

Two details deserve attention. First, the native context window is 256K, not 1 million tokens as widely assumed. The 1.01M figure refers to an expandable mode whose quality tradeoffs Alibaba has not yet documented. Second, the MoE routing is extremely sparse - only 11 of 512 experts fire per token (10 routed plus 1 shared). That ~4% parameter activation ratio (95B of 2.4T) is what makes $2/M input economically viable despite the 2.4T total footprint.

MTP (Multi-Token Prediction) training means compatible inference engines can predict multiple tokens per forward pass, potentially improving throughput. Alibaba has not specified which engines support this.

Self-Hosting Reality: What 2.4T Parameters Actually Costs

The specs sound impressive until you try to load them. Here is the arithmetic.

At 4-bit quantization, 2.4 trillion parameters consume approximately 1.2 TB of storage - before accounting for KV cache, activation memory, or runtime overhead. An NVIDIA H200 GPU has 141 GB of memory. Eight H200s give you 1,128 GB total, which is not enough to hold the weights and serve any meaningful context window simultaneously.

Deployment ScenarioVRAM Needed (est.)HardwarePractical?
Full precision (16-bit)~4.8 TB34+ H200 GPUsDatacenter only
4-bit quantized~1.2 TB9+ H200 GPUsDatacenter only
1.5-bit quantized~450 GB4+ H200 GPUsTight; quality loss uncertain
Qwen3.8-27B (alternative)~27 GB at 4-bit1 consumer GPUYes

The 95B active-parameter count helps inference throughput per token - the model only routes through a fraction of its experts on each forward pass. But active parameters do not reduce the memory needed to store all 512 expert networks. As GEO Toolbox put it: "Sparse activation makes the model cheap to run at scale, not cheap to own."

For local or single-server deployment, the Qwen3.8-27B variant is the realistic option. It shares the Qwen3.8 training lineage but fits on a single consumer GPU at 4-bit. The 2.4T model is a research and datacenter play.

When the API Is the Smarter Path

Since self-hosting a 2.4T MoE model is out of reach for nearly every team, the API is where practical access happens. Alibaba's Model Studio lists Qwen3.8 Max at $2 per million input tokens and $6 per million output tokens - cheaper than both its predecessor and its closest Chinese rival.

API Pricing: Qwen3.8 Max vs Competitors ($/M tokens)
ModelInput ($/M)Output ($/M)Context
Qwen3.8 Max$2.00$6.00256K native (1.01M expandable)
Qwen3.7 Max$2.50$7.501M
Kimi K3$3.00$15.001M

Beyond the per-token API, Alibaba offers three subscription tiers:

  • Lite: $6/month (~10,000 credits)
  • Standard: $18/month (~40,000 credits)
  • Pro: $68/month (~160,000 credits)

The Token Plan endpoint supports both OpenAI-compatible and Anthropic-compatible API formats, so tools like Claude Code, Cursor, and OpenCode work without client-side changes. Qoder, Alibaba's coding product, offers promotional credit rates as low as 0.01x standard during off-peak hours (14:00-00:00 UTC) - but these are launch discounts, not durable pricing.

The subscription model has a sharp edge. Qwen3.8 Max's reasoning mode is verbose, and community reports of quota burn are consistent:

"Qwen tends to overthink A LOT." - u/darko, r/Qwen_AI

In the same Reddit thread, one Lite subscriber reported exhausting their weekly allowance in under two days on roughly 50 million tokens. A Pro-plan user described burning 500 million tokens in a single day across seven parallel agents with a 94% cache-hit rate. The reasoning_effort parameter (settable to low, medium, or xhigh) directly controls how much reasoning output the model generates - and therefore how fast credits disappear. Setting it to low reduced thinking time to roughly 10 seconds per request in one user's test.

For a detailed cost breakdown, see our Qwen3.8 Max API pricing guide.

Benchmarks: Where Qwen3.8 Max Wins and Falls Short

Alibaba published benchmark numbers comparing Qwen3.8 Max against Fable 5 and GPT-5.6 Sol. The picture is mixed - and the gap between vendor-reported and independent scores is large enough to matter.

Qwen3.8 Max vs Fable 5: Alibaba-Reported Benchmark Scores
BenchmarkQwen3.8 Max (Alibaba)Fable 5GPT-5.6 Sol
Terminal-Bench 2.186.684.6Higher than Qwen (exact score not published)
SWE-bench Pro67.780.064.6
GPQA Diamond92.692.6Higher than Qwen
Humanity's Last Exam43.653.3Not published

The independent picture tells a different story. Artificial Analysis measured Qwen3.8 Max at 81.3% on Terminal-Bench 2.1 - a 5.3-point drop from Alibaba's 86.6%. At that independent score, Qwen ranks roughly tenth overall, behind Fable 5 (84.6%) and Kimi K3 (85.0%).

Where Qwen3.8 Max does hold an edge: SWE-bench Pro at 67.7% beats GPT-5.6 Sol's 64.6%, and multimodal scores are strong (MathVision 95.2 vs Fable 5's 92.7, LogicVista 91.9 vs 85.7). The model also improved sharply from Qwen3.7 Max on DeepSWE 1.1 (21.6 -> 56.6), suggesting real gains in software-engineering agent tasks.

Community sentiment reflects the split. The same Reddit thread that complained about overthinking also had this:

"The output was god tier though." - u/TangerineLogical9779, r/Qwen_AI

Speed is reportedly unstable, ranging from 8 tokens/sec to 60 tokens/sec depending on load and time of day. For a head-to-head breakdown against Kimi K3, see our Qwen3.8 Max vs Kimi K3 comparison.

FAQ

Are Qwen3.8-Max weights available to download?

Yes. Alibaba released the Qwen3.8-2.4T-A95B checkpoint on ModelScope on August 12, 2026. This is the first Qwen-Max-tier model with open weights. The cloud-hosted Qwen3.8-Max includes additional production capabilities on top of this base model.

What license do the open weights use?

The exact license has not been confirmed at the time of writing. Historical precedent suggests Apache 2.0 - Qwen3.6 shipped under that license - but Qwen3.7 Max remained closed. Read the actual license file on the ModelScope repository before building anything commercial.

How is Qwen3.8-2.4T-A95B different from Qwen3.8-Max?

Qwen3.8-2.4T-A95B is the open-weight base model: 2.4T parameters, 95B active, MoE architecture, 256K native context. Qwen3.8-Max is Alibaba's cloud version built on this model, augmented with production-environment post-training including unified reward systems and online data balancing. The open weights give you the base; the API gives you the enhanced version.

Can I run Qwen3.8-Max locally?

Not practically. At 4-bit quantization, the model requires approximately 1.2 TB of VRAM - nine or more H200 GPUs for weights alone, with no room for context cache. Even 1.5-bit quantization leaves hundreds of gigabytes. The Qwen3.8-27B variant is the realistic local alternative, fitting on a single consumer GPU.

How much does the Qwen3.8 Max API cost?

The Model Studio API charges $2 per million input tokens and $6 per million output tokens. Subscription tiers start at $6/month (Lite) and go up to $68/month (Pro). Qoder offers promotional credit discounts up to 98% off during off-peak hours, but these are launch pricing and may change.