Two Chinese mixture-of-experts flagships with one-million-token context, released three weeks apart and both reasoning at maximum effort by default, yet priced nothing alike. Kimi K3 is the stronger coder but bills output at $15 per million tokens, 2.5× Qwen 3.8 Max's $6, while Qwen 3.8 Max is the cheaper, fresher pick (GA August 3, 2026) that wins on charts and computer-use. The trade-off: because both models think maximally out of the box, Kimi K3's higher output rate compounds on every request, and Qwen's open weights do not land until about August 11.
Two near-identical flagships, one real split
Qwen 3.8 Max and Kimi K3 are both trillion-parameter mixture-of-experts models with one-million-token context windows and native vision, released within three weeks of each other in mid-2026. On paper their specifications sit in the same tier, which is why most comparisons flatten into benchmark ping-pong. The decision that matters is narrower: how much output cost you will accept for a coding-quality edge, and whether you need open weights or strong multimodal right now.
| Specification | Qwen 3.8 Max | Kimi K3 |
|---|---|---|
| Total parameters | 2.4T | 2.8T |
| Active per token | ~95B | 104B |
| Architecture | MoE (~4% active) | MoE (16 of 896 experts) |
| Context window | ~1M (991K max input) | 1,048,576 |
| Vision | Yes (text + image input) | Yes (native) |
| Default reasoning | xhigh, thinking on | max effort, reasoning on |
| Open weights | No (expected ~Aug 11) | Yes (HF, Jul 27, 1.56TB MXFP4) |
| API status | GA, Aug 3 2026 | GA |
Sources: yottalabs head-to-head, Alibaba QwenCloud via AIReiter, and Moonshot platform.kimi.ai, verified Aug 6, 2026.
Pricing, cost at a workload, and how to access
Both vendors publish per-token pricing, and the gap is not where you would guess from the headline parameter counts. It is concentrated in output tokens, which is where reasoning models spend most of their budget. Kimi K3 charges $15 per million output tokens against Qwen 3.8 Max's $6, and both models think out loud by default, so that output rate applies to reasoning traces as well as final answers.
| Price per 1M tokens | Qwen 3.8 Max | Kimi K3 |
|---|---|---|
| Input (cache miss) | $2.00 | $3.00 |
| Input (cache hit) | $0.25 | $0.30 |
| Output (incl. thinking) | $6.00 | $15.00 |
| Context | ~1M | 1,048,576 |
Sources: Alibaba QwenCloud model page (updated Aug 2, 2026) and Moonshot platform.kimi.ai/docs/pricing/chat-k3, verified Aug 6, 2026.
Run a typical coding workload of 20M input tokens at a 60% cache-hit rate and 5M output tokens, and Kimi K3 lands near $103 ($27.60 input plus $75 output) against about $49 for Qwen 3.8 Max ($19 input plus $30 output). The 2× gap is almost entirely output: caching narrows the input difference, but nothing narrows the output difference, and because both models default to maximum reasoning effort, the output bucket is large.
Access is asymmetric. Qwen 3.8 Max reached general availability on Alibaba's API on August 3, 2026, but its weights are not public; the model page still lists Open Source: No, with weights expected on Hugging Face and ModelScope the week of August 11 (Alibaba, via the Qwen 3.8 Max pricing page). Kimi K3 is the open option: Moonshot posted the 1.56 TB MXFP4 checkpoint on July 27, and the API is live.
Kimi K3 is also reachable through AIReiter's Kimi K3 API endpoint at Moonshot's official rate with no markup, useful if you want to test K3 alongside other models on a single bill. Qwen 3.8 Max is not yet integrated there.
Where Kimi K3 pulls ahead: agentic coding
Kimi K3 is the model to pick when the job is multi-step coding: agent loops, repository refactors, SWE-style tasks where it has to call tools and recover from dead ends. Across the public agentic benchmarks it consistently outpoints Qwen 3.8 Max, and the gap is widest on the tasks that map to real engineering work.
| Benchmark | Qwen 3.8 Max | Kimi K3 |
|---|---|---|
| FrontierSWE | 73.5 | 81.2 |
| TerminalBench 2.1 | 86.6 | 88.3 |
| CharXiv (with tool) | 88.4 | 91.3 |
Source: yottalabs head-to-head, verified Aug 6, 2026.
The community verdict tracks the numbers, with the cost caveat that defines this comparison:
"K3 is god tier. Qwen 3.8 is almost similar to it. However the subscription or API costs are not justified for either of the models as they are…" — r/Qwen_AI
"K3 mogs Qwen 3.8 on this exercise by far, Qwen was 3 times faster than Kimi to generate this landing page." — @filicroval on X
Kimi K3 also carries a higher raw intelligence score, an Artificial Analysis Intelligence Index of 57, but it is slow: roughly 39 output tokens per second and a 3.1-second time to first token (Artificial Analysis). The trade-off is quality and openness in exchange for latency and cost.
Where Qwen 3.8 Max pulls ahead: multimodal and value
Qwen 3.8 Max is the pick when the workload leans on reading charts, screenshots, or driving a browser, and when output cost matters more than squeezing the last few points off a coding benchmark. It beats Kimi K3 on document and perception benchmarks and holds a verified state-of-the-art on one computer-use benchmark, while charging 40% of Kimi K3's output rate.
| Benchmark | Qwen 3.8 Max | Kimi K3 |
|---|---|---|
| CharXiv (no tool) | 93.5 | 84.8 |
| PerceptionBench | 63.5 | 58.5 |
| OSWorld-Verified | 86.1% (SOTA) | no public score |
Source: yottalabs head-to-head, verified Aug 6, 2026.
Qwen 3.8 Max's weak spot is exactly where Kimi K3 is strongest and where it meets the frontier closed models. On SWE-bench Pro it scores 67.7, well behind Claude Fable 5's 80 (Alibaba's published table, vendor-run), so it is not the model for the hardest software-engineering evals. Its API is also three days old at the time of writing (general availability landed August 3, 2026), and its weights are still closed (access details above), so anyone who needs to self-host today has to wait or pick Kimi K3.
How to pick, by workload
The right choice depends less on which model is objectively better and more on what you are calling it for. The matrix below maps the common workloads to a pick and the verified reason behind it.
| If your workload is… | Pick | Why |
|---|---|---|
| Agent loops, repo refactors, SWE tasks (cost not the constraint) | Kimi K3 | FrontierSWE 81.2 vs 73.5, TerminalBench 88.3 vs 86.6 |
| Document/chart parsing, browser or computer-use, cost-sensitive | Qwen 3.8 Max | CharXiv 93.5, OSWorld 86.1% SOTA, $6 vs $15 output |
| Self-host today | Kimi K3 | Weights on HF since Jul 27; Qwen's land ~Aug 11 |
FAQ
Is Qwen 3.8 Max better than Kimi K3 for coding?
No. Kimi K3 leads the agentic coding benchmarks that matter (FrontierSWE 81.2 to 73.5, TerminalBench 2.1 88.3 to 86.6), so for multi-step engineering tasks Kimi K3 is the stronger pick.
Which is cheaper, Qwen 3.8 Max or Kimi K3?
Qwen 3.8 Max. It charges $6 per million output tokens against Kimi K3's $15, and $2 input against $3, with similar cache-hit pricing, roughly a 2× gap at a typical coding workload.
Which is cheaper for agentic coding at maximum reasoning?
Still Qwen 3.8 Max: both models bill reasoning tokens as output at maximum effort, so Kimi K3's $15/M rate widens the gap on exactly the coding workload where it wins.
Do both models support a one-million-token context?
Yes. Qwen 3.8 Max accepts roughly 991K input tokens (less when thinking is on) and Kimi K3 takes 1,048,576; both are nominally one-million-token windows.
Are the weights open?
Kimi K3's are. Moonshot released the 1.56 TB MXFP4 checkpoint on July 27, 2026. Qwen 3.8 Max's are not yet; Alibaba lists them as closed, with a Hugging Face and ModelScope release expected the week of August 11.
Related reading
- Kimi K3 model page: a single-model breakdown of Kimi K3
- Qwen 3.8 Max API pricing: full pricing and technical breakdown for Qwen 3.8 Max
- Qwen3.8-27B: the smaller open-weight Qwen 3.8 release for self-hosting
- Kimi K2.7 Code vs GLM 5.2: an adjacent comparison in the same coding-model cluster