AIREITER

Muse Spark 1.3 API Pricing and Review (2026)

Last Updated: 2026-09-03 02:48:25

Muse Spark 1.3 is currently listed at $1.25 per million input tokens and $4.25 per million output tokens on OpenRouter, but that is not the whole buying decision. It looks especially promising for long-context coding and agent work; the rollout is young, max reasoning was not live at launch, and first-token latency can be high.

What Muse Spark 1.3 is—and what “available” means

Muse Spark 1.3 is Meta’s proprietary reasoning model for long-running agentic work, coding, and multimodal tasks. Meta announced it on September 2, 2026, saying it was rolling out in Muse Code and the Meta Model API; existing reasoning modes were available first, while max reasoning required additional safety testing. (Meta AI Research’s announcement)

Muse Spark 1.3 official Meta announcement

Because the launch graphic includes max results while max reasoning required additional safety testing, buyers should not assume those results are reproducible on day one.

Meta says Muse Spark 1.3 is better at retaining constraints across long tasks, handling interruptions in one thread, asking clarifying questions, requesting confirmation before consequential actions, and acknowledging when it is stuck. It also claims about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 in internal coding comparisons; those are vendor-reported figures without a reproducible task set or error bars.

The launch post describes open weights as future roadmap work. The sources checked here list Muse Spark 1.3 as a hosted/API-only model and do not document a downloadable checkpoint, confirmed license, parameter count, or hardware requirement.

The API price is simple; the cost decision is not

The clearest public rate card is the OpenRouter Muse Spark 1.3 page, not Meta’s launch post. OpenRouter lists standard input and output rates of $1.25 and $4.25 per million tokens, along with separate cache-read and web-search charges.

UsageListed priceWhat to watch
Input$1.25 / 1M tokensFresh prompt and context tokens
Output$4.25 / 1M tokensAbout 3.4× the input rate
Cache read$0.15 / 1M tokensUseful for repeated agent context
Web search$2.50 / 1,000 callsSeparate from token charges
Muse Spark 1.3 API pricing on OpenRouter

A request with 100,000 fresh input tokens and 20,000 output tokens costs about $0.21 before cache or web-search charges: $0.125 for input plus $0.085 for output. Agent loops, retries, and repeated context make complete tasks costlier.

Artificial Analysis reports a $0.78 per million token blended price using a 7:2:1 cache-hit, input, and output mix. Fresh-context or output-heavy workflows will cost more, so use this as a scenario estimate rather than a universal rate.

Do not carry the Contributor assumptions from the Muse Spark 1.2 API pricing guide into Muse Spark 1.3 automatically. The 1.3 public listing establishes the standard rate above, but the sources checked here do not establish a complete 1.3 Contributor contract covering data use, rate limits, or regional availability. Budget against the standard rate until your provider confirms otherwise.

The evidence says “strong coding model,” not “wins everything”

The evidence supports Muse Spark 1.3 as a strong coding and long-context model, not a universal leader. The table below uses scores from Meta’s linked evaluation methodology; benchmark scales differ, and a higher number is comparable only within its own row.

EvaluationMuse 1.3 maxMuse 1.3 xhighMuse 1.2 xhighGPT-5.6 SolClaude Opus 5Highest shown
MRCR 512K–1M98.193.155.573.8—Muse 1.3 max
DeepSWE v1.175.4—55.073.074.0Muse 1.3 max
SWE-Atlas Codebase QnA59.454.046.253.552.7Muse 1.3 max
Terminal-Bench 2.188.889.282.988.886.7Muse 1.3 xhigh
OSWorld 2.066.957.247.662.768.3Claude Opus 5
DeepSearchQA89.489.485.993.090.4GPT-5.6 Sol
GDPVal-AA v217541709161517101824Claude Opus 5

The clearest case for Muse Spark 1.3 is long-context coding: its max result is 98.1 on MRCR at 512K–1M context and 75.4 on DeepSWE, compared with 55.5 and 55.0 for Muse Spark 1.2 xhigh. The same table shows stronger alternatives on some professional-work, computer-use, and browsing evaluations.

Meta’s methodology says the cells can come from Meta’s evaluation, an official leaderboard, or a provider’s self-reported result. Coding tests used restricted tools and no external internet access, so these are useful signals rather than one independently rerun comparison.

Artificial Analysis reports a 61 Intelligence Index score, ranked #10 of 196 comparable models, versus a comparable-model median of 36. ModelCap shows an 81.7 ModelCap Index and #4 board position, but assigns only 20% support and one public benchmark observation. The different systems reinforce one caution: the evidence base is still thin.

For a real workflow, test Muse Spark 1.3 first on repository exploration, long terminal agents, and codebase question-answering. Do not generalize those results to every browsing-heavy or professional computer-use task.

The speed paradox: fast after it starts, slow before it starts

Muse Spark 1.3 appears fast once it is generating, but the first visible answer can take much longer. That favors sustained or batch work over an interactive coding loop where every turn needs to begin immediately.

Artificial Analysis reports 181.7 output tokens per second, versus a comparable-model median of 68.1. It also reports 27.51 seconds to the first token, versus a 3.04-second median, and estimates 41.26 seconds for a 500-token response under its workload. Its provider page shows a 38.51-second time to first answer token and a $0.55 cost per Intelligence Index task.

OpenRouter displays different P50 figures: 62 tokens per second and 1.83 seconds of latency. The metrics are not directly interchangeable because the pages use different workloads and definitions, and Artificial Analysis includes reasoning time in its first-answer metric. Measure both time-to-first-token and total task time in your own harness.

“Since a few peeps asking, my early read (running rn) - feels mechanical (excellent for work, idk for strategy) and very capable. Seems very fast.” — u/NewYak4281, r/singularity

That uncontrolled first read is consistent with the measured trade-off: fast generation after a potentially long wait, plus an interaction style that may feel mechanical. It is useful context, not a substitute for a controlled test.

API capabilities and rollout limits to check before production

Muse Spark 1.3 has the integration features expected from a modern agent model, but several production details remain provider-dependent or undisclosed. The strongest current API evidence comes from OpenRouter’s live model page and Artificial Analysis’s provider tracker, not a complete public Meta specification in the launch post.

Capability or limitCurrent evidenceProduction implication
Context windowOpenRouter and ModelCap list 1,048,576 tokensConfirm the effective endpoint limit
Maximum outputModelCap lists 944K tokensProvider-level output limits may be lower
InputsText, image, video, and files; audio is listed with a warningTreat audio understanding as incomplete
OutputTextPlan a text-mediated workflow
Tool callingListed as supportedTest schemas and tool errors in your SDK
Structured outputJSON-schema response format is listedTest strict-schema failures
Provider coverageOne Meta provider is visible on the OpenRouter page and Artificial Analysis trackerDo not assume failover redundancy
Model weightsAPI-only; no downloadable weights listedNo self-hosting path is established

ModelCap warns that an individual provider can expose a lower context or output limit than the published maximum. The OpenRouter availability panel shows 99.97% uptime over three days but lower availability figures for its 72-hour and 24-hour views. Those are different page metrics, so use current status data and your own error logs when routing production traffic.

Meta does not publish parameter count, architecture, formal API rate limits, licensing terms, or detailed safety-test results in the launch post. It does claim better resistance to prompt injection and better judgment around irreversible actions, without giving attack-success rates or a reproducible red-team protocol.

Should you use Muse Spark 1.3 now?

Muse Spark 1.3 is worth testing now for long-context coding and tool-heavy agents where low token cost and sustained generation matter more than instant first responses. It is not ready to be the only production model for latency-sensitive, audio-heavy, privacy-sensitive, or redundancy-critical systems.

Your situationDecision
Large repository, codebase Q&A, long terminal taskTest Muse Spark 1.3 first
Batch agent work where output speed mattersGood candidate; measure total task cost
Interactive pair programmingUse only if first-token delay is acceptable
Audio-led workflowWait or choose confirmed audio support
Sensitive client code or regulated dataVerify the data contract first
Mission-critical service needing failoverKeep a second model and route

A low-risk rollout is more useful than another leaderboard screenshot:

  1. Start with a non-sensitive repository and five to ten fixed tasks.
  2. Restrict shell, file-write, network, and destructive tools to minimum permissions.
  3. Record completion rate, tool calls, first-token latency, duration, tokens, and invoice cost.
  4. Compare the same tasks with your current baseline and keep a fallback for timeouts or low-confidence edits.

Put Muse Spark 1.3 behind an experiment flag for coding agents, then decide from the repository test rather than the launch-day ranking. The price and long-context results justify the test; slow initial response, limited provider liquidity, incomplete audio support, and provisional benchmark coverage argue against exclusive deployment.

Muse Spark 1.3 API and pricing FAQ

Is Muse Spark 1.3 officially released?

Yes. Meta’s September 2, 2026 announcement says Muse Spark 1.3 was rolling out in Muse Code and the Meta Model API. Max reasoning was still pending additional safety testing.

Where can I access Muse Spark 1.3?

Meta names Muse Code and the Meta Model API. OpenRouter also lists meta/muse-spark-1.3; access can vary by provider, account, and region.

How much does Muse Spark 1.3 cost?

The OpenRouter listing shows $1.25 per million input tokens, $4.25 per million output tokens, $0.15 per million cache-read tokens, and $2.50 per 1,000 web-search calls. Treat these as the cited OpenRouter rates unless your direct provider confirms its own price sheet.

Does Muse Spark 1.3 have a one-million-token context window?

OpenRouter and ModelCap list 1,048,576 tokens, while Artificial Analysis rounds it to 1 million. Confirm the effective context and output limits on the endpoint you call.

Does it support vision, video, tools, and JSON output?

OpenRouter lists text, image, video, and file inputs, plus tool calling and JSON-schema structured outputs. It also lists audio with a warning that understanding is not fully supported.

Is Muse Spark 1.3 open source or self-hostable?

No downloadable weights or self-hosting package were established in the sources checked. Meta describes open weights as future roadmap work, not an available Muse Spark 1.3 artifact.

Is max reasoning available?

Meta said existing reasoning modes were available first and max reasoning would follow additional safety testing. The announcement gave no firm date.

Can I assume Muse Spark 1.2 Contributor pricing applies to 1.3?

No. The 1.2 Contributor terms are documented separately, but the sources checked here do not establish a complete 1.3 Contributor contract. Budget against the standard 1.3 rate until your provider confirms a different tier, data policy, and rate limit.

Run a small, non-sensitive repository benchmark before production: compare the same tasks with your current model while recording first-token delay, completed changes, tool calls, and cost. That will answer the deployment question more reliably than launch-day rankings.