August 2026 changed the Qwen 3.8 vs Kimi K3 question: with both families' weights now downloadable and both flagships in general availability, the answer splits by deployment lane — Kimi K3 holds the strongest verified agentic-coding record, Qwen 3.8 Max wins on token price, and Qwen's Apache-2.0 27B runs on one consumer GPU. The catch on Kimi's side: its "open" flagship is a reported 1.56 TB checkpoint, so openness and practical self-hosting are not the same thing.
What Qwen 3.8 and Kimi K3 actually shipped
Kimi K3 (Moonshot AI) launched July 16, 2026 as a single flagship: a 2.8-trillion-parameter MoE with 104B active per token, Kimi Delta Attention plus Gated MLA, a 1,048,576-token context window, and native vision. Weights and a technical report followed on July 27 under the custom Kimi K3 License — terms summarized in Emergent's comparison attach conditions to large-scale commercial resellers and products above 100 million monthly users.
Alibaba's counter was a family, not a model. Qwen 3.8 Max previewed July 19, reached general availability on QwenCloud August 3 at 2.4T total / ~95B active parameters, and its weights appeared on Hugging Face around August 12 as Qwen/Qwen3.8-2.4T-A95B, under a custom qwen3.8-max license. Days later came Qwen/Qwen3.8-27B: a dense 27B vision-language model under Apache-2.0, native 262K context, extendable to 1M with YaRN (dated August 13–14 by Yotta Labs).
| Kimi K3 | Qwen 3.8 Max | Qwen3.8-27B | |
|---|---|---|---|
| Total / active params | 2.8T / 104B | 2.4T / ~95B | 27B dense |
| Context | 1,048,576 | 1M (991.8K max input) | 262K native, 1M via YaRN |
| License | Kimi K3 License (custom) | Custom qwen3.8-max license | Apache-2.0 |
| Weights live since | July 27, 2026 | ~August 12, 2026 | ~August 13–14, 2026 |
| Vision input | Text + image (formal) | Text, image, video via API | Text, image, video |
Moonshot also reportedly paused new consumer subscriptions within 48 hours of the K3 release as demand outran GPU capacity — worth remembering before betting a product on one provider.
The numbers, sorted by who measured them
Both vendors publish big benchmark tables, and both tables mix harnesses. Kimi's own card flags that competitor scores come from different agent harnesses and, in some rows, H20 rather than H100 hardware. Treat the shared rows below as vendor-reported, not audited.
On those shared rows, Kimi K3 leads where agentic coding lives: per its model card, FrontierSWE 81.2 vs 73.5 and Terminal-Bench 2.1 at 88.3 vs 86.6 (Qwen side: the 2.4T model card). Qwen's launch materials claim OSWorld-Verified 86.1 against the 84.8 on K3's card, and the Qwen card holds single-row standouts — PaperBench 93.0, IFBench 82.8, SWE-bench Pro 67.7 — while K3 posts SWE-Marathon 42.0, BrowseComp 91.2, and MCPMark-Verified 94.5 on rows Qwen doesn't report.
The independent record is lopsided so far. Emergent's tracking credits Kimi K3 with an Artificial Analysis Intelligence Index of 57 at launch and 60 on the current index version, ranked behind Claude Fable 5 and GPT-5.6 Sol, and Orcarouter's comparison adds a #1 Frontend Code Arena placement at 1,679 Elo. Qwen 3.8 Max was not independently indexed at launch; the tracker at cheapestinference.com reported 53 as an Artificial Analysis index figure, and community posts claiming an update to 56 remain unverified.
The only same-task head-to-head, TrilogyAI's StackPerf architecture analysis (documented by Orcarouter), scored K3 83/100 against Qwen's 80/100 on the preview endpoint. Qwen logged 354 repository citations to Kimi's 274, 22 gateway requests to Kimi's 53, and zero failed tool calls to Kimi's two.
Now the row the shared tables skip: on Qwen's own card, the 27B posts OSWorld-Verified 84.3 against K3's 84.8. A 27B dense model sits within half a point of a 2.8T flagship on computer use. It is not in K3's coding class (Terminal-Bench 73.0 vs 88.3), but LiveCodeBench v6 at 90.3 and SWE-bench Pro at 61.7 make it a real workhorse. A hands-on comparison thread on r/AISEOInsider split the same way the data does: Qwen winning more one-shot builds, Kimi pulling ahead once context and revision depth grew.
API cost: the $15 output line and the reasoning tax
List prices point one direction. Real spend points to the same place, harder, because both models think by default — K3 at reasoning_effort: max, Qwen at xhigh — and reasoning tokens bill as output.
| Per 1M tokens | Qwen 3.8 Max (QwenCloud) | Kimi K3 (Moonshot) |
|---|---|---|
| Input (cache miss) | $2.00 | $3.00 |
| Input (cache hit) | $0.25 | $0.30 |
| Output, incl. reasoning | $6.00 | $15.00 |
| Explicit cache create / read | $2.50 / $0.17 | — |
Two worked sessions at August 2026 list prices:
| Session | Qwen 3.8 Max | Kimi K3 | Gap |
|---|---|---|---|
| Agentic coding: 500K input (60% cached), 40K output | $0.72 | $1.29 | 1.8× |
| Fresh-context doc pipeline: 2M input, 100K output | $4.60 | $7.50 | 1.6× |
A routing rule falls out of that math: send a task to K3 only when its acceptance rate on your workload beats Qwen's by more than the ~1.8× token-cost gap. Moonshot itself sells the budget alternative — Kimi K2.7 Code at $0.95 input / $4.00 output over 256K context — so the cheapest Kimi-family coding route is not K3. If you're routing K3 today, its API endpoint is listed here at Moonshot's published rates, and the deeper price mechanics for both sides are broken out in the Qwen3.8-Max pricing analysis and the Kimi K3 pricing guide.
Self-hosting: 1.56 TB versus one RTX 4090
This is where the two open-weight stories diverge by roughly two orders of magnitude.
K3 ships quantization-aware MXFP4 weights (model card), but a checkpoint reported at 1.56 TB is a cluster project — the same source recommends vLLM across 64 or more accelerators. That's a rentable-but-real line item before you've served a single token, and the license's large-scale conditions apply on top.
The 27B is the opposite story. Community GGUF builds span roughly 8.5 GB to 28.9 GB, Unsloth's dynamic GGUF and NVFP4 builds slot into about 17 GB minimum, and the official FP8 variant exists for server deployment. It runs on hardware people already own:
"Qwen3.8-27B at 160K context on a SINGLE RTX 4090 — 47–57 tok/s, full GPU offload" — r/Qwen_AI deployment post
The Max weights sit in between: Qwen3.8-2.4T-A95B and its FP8 sibling are downloadable under the custom qwen3.8-max license, but the open artifact is text-only with thinking that cannot be disabled — vision input, a non-thinking mode, and 1M default context live in the QwenCloud API layer, not the checkpoint. For everything the 27B can and can't do at each quant, the Qwen3.8-27B field guide has the runtimes and VRAM tables; K3's weight release is detailed in the Kimi K3 open-weights writeup.
Integration details that decide agent projects
Three card-level behaviors will cost you a debugging day if you meet them in production.
| Model | Behavior / limit | Implementation consequence |
|---|---|---|
| Kimi K3 | Thinking always on; multi-turn and tool-use clients must resend the full prior assistant message, including reasoning_content and tool_calls (card) | Drop the reasoning trace and agent loops degrade; keep it and spend grows with loop length |
| Qwen3.8-27B | preserve_thinking on by default; reasoning_effort drops to medium/low; YaRN past 262K is static scaling (card) | Disabling per request is supported; the card warns lower effort doesn't always cut end-to-end time because retries eat the savings; enable YaRN only for long jobs |
| Qwen 3.8 Max vs K3 caps | Max: 991.8K input / 131.07K output (QwenCloud); K3: 1,048,576 input / 131K default output rising toward 1M | Neither "1M context" promises 1M useful output; a third-party needle test found 18/18 facts at 248K and tested nothing beyond |
FAQ
Is Qwen 3.8 better than Kimi K3 for coding?
On the evidence available, no. Kimi K3 leads the agentic-coding rows that matter (FrontierSWE 81.2 vs 73.5, Terminal-Bench 2.1 88.3 vs 86.6) and holds the only independent index scores. Qwen 3.8 Max is close on terminal work and cheaper per token and in the modeled sessions; the 27B is a different weight class entirely.
Which is cheaper, Qwen 3.8 or Kimi K3?
Qwen 3.8 Max on every list line: $2 vs $3 input, $6 vs $15 output. Because both models bill default-on reasoning as output, Kimi's disadvantage grows with task difficulty — roughly 1.8× on a cached agentic session in the math above.
Can you self-host them, and under what licenses?
Yes, all three checkpoints are downloadable. Qwen3.8-27B is Apache-2.0 and fits a single 24 GB GPU quantized; Kimi K3 needs cluster-scale hardware for its ~1.56 TB MXFP4 checkpoint and carries the custom Kimi K3 License; Qwen3.8-Max weights are public as Qwen3.8-2.4T-A95B under the custom qwen3.8-max license.
Is the 1M context window real on both?
Roughly. Kimi K3 specifies 1,048,576 tokens; Qwen 3.8 Max lists 1M with 991.8K maximum input; the 27B is native to 262,144 and reaches 1M only through YaRN scaling that trades short-context quality. Independent verification exists only at the 248K level.
The pick, and what would change it
| Your lane | Pick | Why |
|---|---|---|
| API coding agents, quality first | Kimi K3 | FrontierSWE 81.2, Terminal-Bench 88.3, AA-verified 60 |
| API work where spend matters, docs/vision heavy | Qwen 3.8 Max | $6 vs $15 output, OSWorld 86.1, explicit cache at $0.17/read |
| Self-host on hardware you own | Qwen3.8-27B | Apache-2.0, ~17–29 GB quants, 160K context on one RTX 4090 |
| Self-host a frontier flagship | Kimi K3 | The one open checkpoint here with independent frontier-class scores — budget a cluster |
Two things would flip parts of this table: independent scoring of Qwen 3.8 Max (tracker-reported 53 vs K3's verified 60, built on thin evidence) and published terms for the custom qwen3.8-max license. Until either lands, the Qwen 3.8 Max vs Kimi K3 API head-to-head covers the API-only lane in more depth.
Related reading: Qwen3.8-27B: What's Confirmed and What 17GB Really Runs · Kimi K3 Open Weights · Qwen 3.8 Max vs Kimi K3: API Edition