Qwen3.6-27B shipped in April 2026 with a full benchmark table on its Hugging Face model card. Muse Glimmer 30B followed in August 2026 as a Meta open-weights release, and within days the local-LLM community was running head-to-head tests. The headline numbers point one direction: Qwen scores 60.7 on TerminalBench 2.1 versus Glimmer's 51.7, and Qwen's official SWE-bench Verified sits at 77.2, but the full picture includes a VRAM advantage for Glimmer that changes which model you can run on a single GPU.
TerminalBench 2.1: Qwen Leads by 9 Points
TerminalBench measures multi-step terminal-agent reliability, including tool use and sustained context over extended sessions. The linked community discussions frequently reference TerminalBench 2.1 as the available head-to-head coding-agent result.
Community-reported TerminalBench 2.1 scores, surfaced in r/LocalLLaMA discussions from August 10-11, 2026:
| Model | TerminalBench 2.1 | Source type |
|---|---|---|
| Qwen3.6-27B | 60.7 | Community-reported |
| Muse Glimmer 30B | 51.7 | Community-reported |
| Gemma 4 31B | 43.4 | Community-reported |
One caveat: TerminalBench evaluates the model-plus-harness combination, not the base model alone. Qwen3.6-27B's official card specifies a Harbor/Terminus-2 harness with a 3-hour timeout, 32 CPUs, 48 GB RAM, 80K max output, 256K context, and five-run averaging. Glimmer's score may reflect a different or less-optimized harness configuration, which means the 9-point delta could narrow or widen depending on your agent setup.
Published Coding Benchmarks: Qwen Has Results; Glimmer Does Not Yet
Qwen3.6-27B has official benchmark scores published on its model card. In the linked sources surveyed for this comparison, no official SWE-bench, LiveCodeBench, or comparable agentic coding benchmark was found for Muse Glimmer as of August 2026. This means Qwen has the stronger published evidence base, but a like-for-like official head-to-head does not exist.
Qwen's official coding benchmarks:
| Benchmark | Qwen3.6-27B (official) |
|---|---|
| SWE-bench Verified | 77.2 |
| SWE-bench Pro | 53.5 |
| SWE-bench Multilingual | 71.3 |
| Terminal-Bench 2.0 | 59.3 |
| LiveCodeBench v6 | 83.9 |
These are vendor-published results. The model card notes that SWE-bench evaluations use Qwen's internal bash/file-edit scaffold at temperature 1.0, top-p 0.95, with a 200K context window. SWE-bench Pro results were computed on a refined task set where Qwen corrected problematic items, so direct comparison to public leaderboard values may not be apples-to-apples. Third-party coverage from morphllm flags that these results rely on Qwen's own agent scaffold with limited independent reproduction.
Real-World Coding Tests: What Local Testers Found
OpenCode Q4 Test on M5 Pro
A developer ran Muse Glimmer at Q4 quantization (Unsloth build) on an M5 Pro with 48 GB RAM through OpenCode. The model used approximately 20 GB RAM and generated at 17 tokens/second. The verdict:
"Overall, sits below Qwen3.6 27B" - u/curiousily_
Frontend and backend coding output was rated below Qwen, though the author noted a positive: Muse Glimmer did not fail any tool calls during the test. No reasoning loops or extended thinking was configured, which may have limited Glimmer's performance.
The max_tokens Gotcha
A separate tester on r/LocalLLM discovered that Muse Glimmer can appear much worse than it is when the output token budget is too small. The model consumes its token budget in reasoning before producing visible output, resulting in empty or truncated answers. Raising max_tokens improved their test harness from 6/13 passing tasks to 11/13, nearly doubling the pass rate without changing the model itself. If you test Glimmer with default output limits, you may be measuring your configuration, not the model.
Hermes Terminal Looping
Another user reported Muse Glimmer getting stuck with excessive terminal commands when paired with the Hermes agent framework, a behavior they had not encountered with Qwen3.6-27B on the same setup. This anecdote is directionally consistent with the available TerminalBench gap, though it does not isolate the model from the Hermes configuration.
Token Efficiency: A Qualitative Lead
The benchmark gallery post by u/NoFaithlessness951 raised an efficiency angle:
"Little less smart than Qwen, but way fewer tokens per task." - u/NoFaithlessness951
This is a qualitative community observation, not a controlled token-per-successful-task measurement. No head-to-head study has quantified Glimmer's token advantage over Qwen with matched tasks, success rates, and latency data. Treat the efficiency claim as a promising lead rather than an established fact.
VRAM and Context: Glimmer's Hardware Advantage
This is where Muse Glimmer pulls ahead decisively. A tester on r/LocalLLaMA demonstrated that Muse Glimmer 30B at Q4_K_XL with DFlash speculative decoding, multimodal projector, and full F16 KV cache fits within approximately 22-23 GB on a single RTX 3090, with a 262,144-token context window configured.
Same RTX 3090, same Q4_K_XL quantization:
| Model | F16 KV Context | Q8 KV Context | VRAM Used |
|---|---|---|---|
| Muse Glimmer 30B | 262,144 | N/A | ~22-23 GB |
| Qwen3.6-27B | 70,000 | 125,000 | Fits 24 GB |
| Gemma 4 31B | 52,000 | 81,000 | Fits 24 GB |
The author called Qwen's 70K F16 context "borderline unusable" for workflows that load large repository context into the prompt.
Reported Muse Glimmer throughput on that RTX 3090 setup:
- Generation: 64-124 tokens/second (varying by code vs prose)
- Prompt processing: ~1,400 tokens/second
- Long-context retrieval: passed a two-needle haystack test at ~150K tokens on the first attempt
Qwen3.6-27B on the same hardware requires either Q8 KV cache compression (trading output fidelity for context length) or offloading to a more powerful machine. The tester reported moving Qwen to their DGX Spark when full context was needed, a significantly more expensive device.
On higher-end hardware, a Qwen3.6-27B NVFP4 build on a DGX Spark reached 28-33 tokens/second single-session across contexts up to 128K, according to kie.ai's benchmark analysis. An Unsloth Muse Glimmer Q5_K_M build on an RTX 5090 reportedly hit 220-253 tokens/second for code-patch generation via a patched llama.cpp DFlash path.
Which Model Should You Run?
| Your scenario | Pick | Why |
|---|---|---|
| Coding-only, one-shot generation | Qwen3.6-27B | Stronger published coding benchmark evidence; leads in the available TerminalBench 2.1 community comparison |
| Large-context agentic coding on a single 24 GB GPU | Muse Glimmer 30B | 262K F16 context vs Qwen's 70K on same hardware (Q4_K_XL/DFlash setup) |
| Long-horizon terminal tasks | Qwen3.6-27B | 60.7 vs 51.7 on TerminalBench 2.1; Glimmer community report of agent-framework looping |
| High-volume, cost-sensitive work | Muse Glimmer 30B (possible) | One community report suggests fewer tokens per task; no matched cost-per-success comparison |
| Repository-level coding with large context | Muse Glimmer 30B (if VRAM-limited) | Qwen needs Q8 KV compression or multi-GPU for large context on 24 GB |
| Maximum coding accuracy, hardware not a constraint | Qwen3.6-27B | Stronger published benchmarks, official scores |
FAQ
Does Muse Glimmer have an official SWE-bench score?
No official SWE-bench, LiveCodeBench, or TerminalBench score for Muse Glimmer 30B was found in the sources surveyed for this comparison as of August 2026. The TerminalBench 2.1 score of 51.7 comes from community testing in Reddit threads, not an official Meta evaluation.
Can either model run on a single RTX 3090?
Yes. Muse Glimmer 30B at Q4_K_XL fits with 262K context and full F16 KV cache in ~22-23 GB. Qwen3.6-27B at Q4_K_XL fits but is limited to 70K F16 context or 125K with Q8 KV cache compression on the same GPU.
Is Muse Glimmer censored for coding tasks?
One user reported Muse Glimmer refusing to help debug Python mouse-control code, framing it as a potential security concern. This is a single anecdotal report with an unspecified system prompt and configuration. It is not sufficient to determine whether the behavior comes from the model itself or the inference setup.
Which is better for agentic coding workflows?
Qwen3.6-27B is more reliable for long-horizon terminal agent tasks based on TerminalBench scores and community reports. Muse Glimmer 30B is more viable on constrained hardware due to its VRAM efficiency and may produce fewer tokens per task, which could reduce cost and latency in high-volume agent loops, though no controlled comparison has quantified this yet.
What max_tokens should I use for Muse Glimmer coding?
Set a generous output budget. A tester found that tight max_tokens limits caused Glimmer to exhaust its budget on reasoning before producing visible output, yielding empty or truncated answers that looked like failures. Raising the limit improved their test harness from 6/13 to 11/13 passing tasks. The Qwen3.6-27B model card recommends 32,768 tokens for general queries and 81,920 for hard benchmark tasks; the same output budget is a reasonable starting point for Glimmer until Meta publishes its own guidance.