AIREITER

Mistral Large 4 API Pricing: What the Preview Really Offers

Last Updated: 2026-10-06 18:50:19

A million-token context and a $0.68 input rate make Mistral Large 4 look unusually cheap for a trillion-parameter model. The catch is that the model is still a public preview: the discounted price has no clear end date, gateway limits differ from the headline specification, and the future open-weight license is not published yet.

The decision in one minute

Mistral Large 4 is worth testing now if you need a low-cost, long-context API for documents, code repositories, or tool-using workflows. It is too early to treat the preview as a stable self-hosting replacement or as a proven leader on coding benchmarks.

NeedDecision
Test long documents or large codebases through an APITry it now, while measuring cache hits and output cost
Need a predictable long-term priceWait or budget for the higher rate of $1.36 input and $4.18 output
Need private deploymentWait for the weights, license, quantization, and serving guidance
Need the best coding scoreDo not choose from the launch numbers alone; run the same harness on your tasks

This is a public-preview API decision, not a final model verdict. Mistral announced the preview on October 6, 2026, and its public materials say weights are planned for the end of October. (Mistral announcement)

What is actually available on October 6

Mistral Large 4 is available as a public-preview API through Mistral Studio. Mistral describes it as a natively multimodal mixture-of-experts model with 1.05 trillion total parameters and 49 billion active parameters per token. The official announcement also describes function calling, structured outputs, document question answering, batching, agents, and built-in tools.

The model is also listed as mistralai/mistral-large-4-0 on OpenRouter, so OpenRouter access is no longer just a launch-day rumor. Limits, latency, billing, and fallback policy can differ by route.

The API is live, but the downloadable weights and final license are not. “Open weight” should not be treated as unrestricted commercial open source until Mistral publishes the model card and terms.

The price is two different numbers for a reason

The current preview rates shown in Mistral's model documentation are attractive:

Token typePreview ratePossible post-preview rate shown in launch materials
Input$0.68 / 1M tokens$1.36 / 1M tokens
Cached input$0.07 / 1M tokens$0.14 / 1M tokens
Output$2.09 / 1M tokens$4.18 / 1M tokens

The lower figures are presented as 50% preview pricing in Vercel's gateway listing and other model-access pages, but no end date is stated. Treat $0.68 and $2.09 as the current rate, not a permanent contract.

For a workload with 500,000 uncached input tokens and 100,000 output tokens, the preview bill is about $0.55: $0.34 for input plus $0.209 for output. If the input is cached, it is about $0.244: $0.035 cached input plus $0.209 output. Evaluate cache reuse, output volume, and the discount’s durability before migrating production traffic.

Context and capability depend on the route

Mistral’s model documentation describes a 1-million-token context window. That is a model-level claim, not necessarily the limit exposed by every gateway.

Vercel’s AI Gateway page currently lists a 524K shared context window and a 262K maximum output for its Mistral provider. It also reports 0.9-second P50 time to first token, 38 tokens per second, and a 74% cache-hit figure based on gateway traffic.

Mistral Large 4 gateway pricing and provider metrics

Those are route-level numbers, not universal model guarantees. The page documents image input, tool calling, OpenAI-compatible Chat Completions and Responses APIs, and Anthropic-compatible Messages access; its gateway ID is mistral/mistral-large-4, unlike OpenRouter’s provider-qualified ID.

For an integration, verify the endpoint’s shared context, output cap, image/token accounting, tool behavior, cache policy, provider routing, and fallback rules. A 1M headline is useful for shortlist decisions, but it is not enough to size an application.

What the benchmark evidence can support

TechsCurrent’s launch analysis reports Mistral’s figures of 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4, and 59.9% on AutomationBench. CellCog’s launch comparison cites DeepSeek V4.1 Flash at 74.2 on DeepSWE and Mistral Large 4 at 61.7, while Mistral’s AutomationBench figure is higher than its comparison rows. These are vendor-reported or cross-source signals, not one standardized leaderboard.

Before using the model for production agents, run a fixed evaluation:

  1. Use the same repository, prompt, tools, and timeout for Mistral Large 4 and the incumbent model.
  2. Measure task completion, tool-argument errors, retries, latency, output tokens, and cost.
  3. Add long-context retrieval cases where the answer appears in the middle and near the end of the input.
  4. Repeat the test with cached and uncached context.
  5. Keep preview results separate from future weight-release results.

Who should use it now

API-first teams with large recurring context: Test Mistral Large 4 now. The $0.07 cached-input rate can matter when an agent repeatedly sends the same policy, repository, or document prefix. Put a spend cap on the experiment because the higher rate remains plausible.

Teams building image-aware document workflows: Test image input and document extraction at the route you plan to deploy. The model supports image input, but route-specific file limits and billing details still need verification.

Coding-agent teams: Use it as a candidate, not an automatic winner. The reported benchmark evidence is mixed, and real reliability depends on tool calls, retries, and completion rate.

Privacy-sensitive or self-hosting teams: Wait for the weights and license. The 49-billion active-parameter figure does not mean the full checkpoint is small; the total model still has to be stored and served. Hardware, quantization, throughput, and redistribution rights are unresolved.

Mistral Large 4 API FAQ

Is Mistral Large 4 open source yet?

No final license or downloadable weights were available at launch. Mistral planned a weight release by the end of October 2026, but “open weight” alone does not define commercial or redistribution rights.

Is the $0.68 price permanent?

No end date is stated. Budget tests at the current $0.68 input and $2.09 output rates, while modeling $1.36 and $4.18.

Is the context window 1M or 524K?

Mistral documents a 1M-token model context; Vercel lists 524K shared context for its gateway route. Confirm the limit for the provider and API format you will call.

Is it already on OpenRouter?

Yes, OpenRouter currently lists mistralai/mistral-large-4-0. Check the live page before deployment because preview routing can change.

Can it handle images and tools?

Official and gateway documentation describe image input and tool calling. Test the request format, image limits, token accounting, and malformed-tool-call behavior in your chosen route.

Should I self-host it when weights arrive?

Not automatically. Wait for the license, checkpoint formats, quantization options, memory requirements, throughput data, and a reference serving setup.

The next test to run

Create a 20-task evaluation before routing real traffic: five long-context retrieval tasks, five coding changes, five image/document questions, and five tool-calling tasks. Log success rate, retries, time to first token, total latency, input/cache/output tokens, and effective cost per completed task.

Sources: Mistral model documentation, Mistral announcement, Vercel AI Gateway, OpenRouter model page, TechsCurrent, CellCog.