The bottom line on Muse Spark 1.2
Meta shipped Muse Spark 1.2 on August 5, 2026 with better coding benchmarks than 1.1 and a new terminal agent built to run it. It still lost all three coding charts to Claude Opus 5. That tension is the whole story. The update is real: Meta scaled coding training compute, co-trained the model with its Muse Code agent, and the numbers moved. But the version-over-version jump is partly the new harness rather than the model alone, and on Terminal-Bench, DeepSWE, and Meta's own internal coding eval, Anthropic's model still leads by 4 to 9 points.
The price is what makes it interesting anyway. Standard pricing held at $1.25/$4.25 per million tokens, and a contributor tier drops that roughly 12x further. The catch: the default path feeds your code into Meta's training, and at the highest reasoning effort the model's time-to-first-token regressed about ninefold. If you are evaluating Muse Spark 1.2 for production coding, those two details matter more than the leaderboard rank.
What actually changed from 1.1
Muse Spark 1.2 is a coding-focused update to Meta's frontier reasoning model, released August 5 alongside Muse Code, a terminal-based coding agent now in beta. Meta itself calls it a "moderate improvement" over Muse Spark 1.1. That is deliberately modest language for an update that concentrates gains where coding agents work hardest: multi-file refactors, long debugging sessions, and tasks that run well past a single prompt.
Two training details drive the bump. First, Muse Code was in the training loop from day one. Meta co-trained the model with rejection-sampled harness trajectories, tuning it to perform best inside its own agent. The company says it trained across multiple harnesses, so the model still generalizes to other coding tools. Second, Meta ran a self-improvement loop: Muse Spark 1.1 generated challenging coding environments and instruction-following templates, then graded candidate solutions against them, producing the training set for 1.2. Meta credits that loop with making 1.2 measurably better at following complex instructions.
The update targets the Muse family's weakest flank: when Muse Spark debuted in April 2026 it trailed on agentic coding, scoring 77.4 on SWE-Bench Verified against Claude Opus 4.6's 80.8. A coding-specialized checkpoint paired with a purpose-built harness is the direct answer, while general agentic capability is maintained.
The benchmarks, read honestly
On Terminal-Bench 2.1, Muse Spark 1.2 running in Muse Code scored 82.9%, edging GPT-5.6 Terra in Codex (81.8%) and Grok 4.5 in Grok Build (81.6%) but trailing Claude Opus 5 at max effort in Claude Code, which leads at 86.7%. On DeepSWE 1.1 it posted 59.3%, third behind Opus 5 (65.0%) and GPT-5.6 Terra (64.8%). Meta's own internal coding benchmark is the most candid of the three: Muse Spark 1.2's 70.6% beats GPT-5.6 Terra (65.4%) and Gemini 3.6 Flash (63.9%) yet still sits nearly nine points behind Opus 5's 79.4%. Claude tops all three charts.
The generational gains are real: Muse Spark 1.2 improves on 1.1 by 6.7 points on Terminal-Bench and 6.3 on DeepSWE. A detail in the chart labels matters for anyone comparing versions, though. The 1.1 scores were recorded in a generic mini-swe-agent harness, while 1.2 ran in Muse Code. Some of that jump belongs to the new harness, not just the new model. On Artificial Analysis's blended Intelligence Index, 1.2 scores 54, ranked #14 of 186 models against a class median of 32. That is up from 1.1's 51, a quieter gain than the coding charts suggest.
The head-to-head against Claude and GPT is exactly the question buyers keep asking. On Reddit, one developer put it plainly before any 1.2 data existed:
"Has anyone used Meta's MUSE Spark 1.1? I'm curious how it compares with Claude and GPT in reasoning, coding, speed, accuracy, context handling."
The 1.2 benchmarks are the first rigorous answer to that question: competitive, clearly second on coding, ahead of the GPT-5.6 and Gemini options on most charts.
Two catches that shape production use
Two practical details determine whether Muse Spark 1.2 works in a real workflow, and neither shows up in a leaderboard rank.
The xhigh latency regression
Muse Spark is a reasoning model. It thinks before it answers, and you cannot disable that. The /effort dial spans five levels from minimal to xhigh, and thinking tokens are billed as output. At the top end the cost is steep in two senses. Independent measurements published by OrcaRouter show time-to-first-token jumping from 1.1's 2.90 seconds to 26.12 seconds at xhigh, roughly ninefold, while output speed fell from 213.5 to 165.0 tokens per second. Running the Artificial Analysis Intelligence Index evaluation on 1.2 cost $637.85, up 16% from 1.1's $548.07, purely because the model deliberates more. For interactive coding where a developer is waiting on a reply, maximum effort trades latency for accuracy in a way that often is not worth it.
The contributor tier trades your code for a discount
Meta prices Muse Spark 1.2 in two tiers, and the one it steers new users toward carries a condition. The contributor tier ($0.10/$0.20 per million input/output tokens, against $1.25/$4.25 on standard) is roughly 12x cheaper on input and 21x on output, with cached input near-free at $0.002. The trade is explicit permission for Meta to use your prompts and completions to train future models. It is the cheapest rate in the comparison above; you pay for it with your data.
The default Muse Code on-ramp lands developers on this tier. In testing reported by VentureBeat, the one-line installer worked on a Mac mini (a 97 MB download and a sign-in), but the agent stopped short of running anything, reporting no visible models and requiring a payment method even on the discounted tier. Low-cost is accurate; free is not. Rate limits are tighter too, at 60 requests per minute versus 3,000 on standard, a clear signal the tier is aimed at individuals and prototyping rather than production. Teams with proprietary codebases need to consciously move to standard pricing or request zero data retention through Meta sales.
Pricing in context: where Muse Spark sits
Muse Spark 1.2's standard tier ($1.25 per million input, $0.15 cached, $4.25 output, 1M-token context window, no long-context premium) sits in the lower-middle of the coding-model market. It is markedly cheaper than the frontier coding leaders: Claude Opus 5 at $5/$25 and GPT-5.5 at $5/$30 run roughly 4 to 7x the output price. It is close to a near-exact price twin, Z.ai's GLM-5.2 at $1.40/$4.40, and a step below Grok 4.5 at $2/$6. The cheap end of the market, DeepSeek V4 Pro at $0.435/$0.87, undercuts it several times over, as does the Muse Spark contributor tier itself.
| Model | Input $/M | Output $/M | Cached $/M | Context |
|---|---|---|---|---|
| Muse Spark 1.2 (standard) | 1.25 | 4.25 | 0.15 | 1M |
| Muse Spark 1.2 (contributor)* | 0.10 | 0.20 | 0.002 | 1M |
| Claude Opus 5 | 5.00 | 25.00 | — | — |
| GPT-5.5 | 5.00 | 30.00 | — | — |
| GLM-5.2 | 1.40 | 4.40 | — | — |
| Grok 4.5 | 2.00 | 6.00 | — | — |
| DeepSeek V4 Pro | 0.435 | 0.87 | — | — |
\*contributor tier grants Meta rights to train on your data; 60 req/min cap.
Where you can actually get it (and what to use today)
Muse Spark 1.2 is available in three places: the Meta Model API, the Muse Code terminal agent, and OpenRouter. It is not yet on every aggregator — AIReiter's catalog, checked August 6, 2026, does not list it. If your stack runs on a unified API and you want a coding model today, you choose among what is already routed.
The closest price-tier match available now is GLM-5.2 at $1.40/$4.40. On August 6, 2026 I ran a one-shot coding-reasoning task through glm-5.2 (routed via the platform's API) to see what "available today" delivers: a Python withdrawal function missing its overdraft check. GLM-5.2 returned in 3.2 seconds (85 prompt, 128 completion tokens), correctly identified the missing balance guard, produced the fix (if amount > 0 and account_balance >= amount:), and wrote a test that fails before the fix and passes after. That is one task, not a benchmark, but it is a working coding model at the same price point, callable now. DeepSeek V4 Pro, Kimi K3, Claude Sonnet 5, and the GPT-5.6 family round out the coding-capable options already in AIReiter's catalog.
Should you use Muse Spark 1.2?
The decision splits three ways depending on what you are optimizing for.
Pick Muse Spark 1.2 if you are doing price-sensitive coding at scale (large multi-file refactors, long debugging sessions, long-horizon agentic tasks) and you are willing to go through Meta, Muse Code, or OpenRouter directly. At $1.25/$4.25 it delivers frontier-class coding capability at mid-tier pricing, and the 1M context window with context compaction suits work that outlasts a single prompt. Keep reasoning effort below xhigh unless your task can absorb a 26-second first token.
Pick Claude Opus 5 if you need the top of the coding charts. It leads all three benchmarks Muse Spark 1.2 appears on. You pay 4 to 6x more per token for the accuracy lead.
Pick a model already in your aggregator if you need to ship today and Muse Spark is not on your platform. GLM-5.2 gives you the same price tier with no wait; DeepSeek V4 Pro is cheaper still where raw cost dominates.
FAQ
Is Muse Spark 1.2 better than Claude Opus 5 at coding?
No. Claude Opus 5 leads on all three coding benchmarks Meta published (Terminal-Bench 2.1, DeepSWE 1.1, and Meta's internal coding eval), by margins of 3.8 to 9 points. Muse Spark 1.2 is a clear second, ahead of GPT-5.6 Terra and Gemini 3.6 Flash on most charts.
Is Muse Spark 1.2 open source?
No. Unlike Meta's Llama family, Muse Spark is proprietary: cloud-only, no downloadable weights, no self-hosting. Mark Zuckerberg has said he will "have more to share" on possible open-source releases, but 1.2 ships with no weights and no license.