On FinanceBench (368 SEC filings, roughly 53,900 pages, 150 questions), Mistral says its new Agentic Search layer lifted answer accuracy by 47 to 53 percentage points over one-shot RAG with a search-only loop, and up to 59 points with the full navigation loop, while cutting token use by up to a third. The catch sits in the latency column: the full loop still needs 154 seconds at p90 and 71 seconds on average.
Mistral Agentic Search, announced August 20, 2026, is a retrieval layer for hard enterprise corpora such as filings, contracts, and scanned PDFs. It is not a web search upgrade, and the minutes-long loop is the price of the accuracy.
What Mistral Agentic Search Is
Mistral Agentic Search is a retrieval layer released on August 20, 2026, that wraps an existing index in a multi-step loop of five tools: search, open, navigate, read, and grep. It is available through the Mistral Search Toolkit and built into Libraries in both Studio and Vibe. The index stays yours, with no fine-tuning required; the agent loop handles the questions that chunk retrieval alone cannot answer.
The Five-Tool Loop: search, open, navigate, read, grep
The mechanics are best understood through Mistral's own worked example: calendar-year 1953 U.S. national defense spending. A single search call surfaced January–June figures but missed July–December. The agentic path instead ran two searches, identified treasury_bulletin_1954_02.pdf, then read page 15, Table 3, and returned all twelve monthly values, totaling 44,463 million dollars for the year.
| Tool | What it does |
|---|---|
search | Queries the index for likely sources |
open | Loads a located document |
navigate | Moves to pages, tables, or sections inside it |
read | Extracts the content at that location |
grep | Pattern-matches exact tokens, codes, or identifiers |
One-shot RAG retrieves a fixed set of chunks and must answer in a single pass; it cannot request another document after weak retrieval. The loop can compare sources, follow references into footnotes and tables, and recover from partial results. Mistral claims the trade is fewer wasted retries and turns, not more.
What the Benchmarks Actually Show
The announcement carries two benchmark suites, both vendor-reported on out-of-the-box Search Toolkit defaults with no tuning; Mistral's own framing is "these results are floors, not ceilings." FinanceBench (Islam et al., 2023; open release on GitHub) uses 368 SEC 10-K, 10-Q, and 8-K filings, about 147 pages each by Mistral's count, with 150 questions that Mistral's setup judges using an LLM calibrated against human labels.
Adding the navigation tools (open, navigate, read, grep) on top of the search-only loop contributes a further +8.7 pp for Mistral Medium 3.5 and +6.7 pp for GLM-5.2. Against the search-only loop, the full navigation loop used 23.9% fewer tokens with Mistral Medium 3.5 and 33.7% fewer with GLM-5.2. Latency moves the same direction: p90 falls from 255 to 154 seconds, and mean from 108 to 71 seconds. Note the comparison target: the token reduction is measured against the search-only agentic loop, not against one-shot RAG.
OfficeQA Pro is the harder suite: 696 historical U.S. Treasury Bulletins, roughly 89,000 pages of scanned, table-heavy government PDFs, 133 questions. GLM-5.2 reached 51.9% accuracy with the full loop, a +45.6 pp jump that implies a 6.3% one-shot baseline; Mistral Medium 3.5 improved +27.1 pp. Navigation added up to 7.0% fewer turns.
Mistral benchmarked with its own Mistral Medium 3.5 and third-party Z.ai GLM-5.2, and both gained in the same pattern. The announcement also cites Kimi research putting GLM-5.2 at 41.4% on OfficeQA Pro under a Claude Code harness versus 51.9% under Mistral's harness; no primary Kimi source is linked, so that 10.5 pp harness gap is Mistral's characterization. Mistral's takeaway, "retrieval quality scales with model capability," says the harness is the bottleneck, not the model.
When One-Shot RAG Still Wins
Mistral itself says indexed retrieval remains the foundation and the agentic loop is for when initial results are insufficient. The decision is mostly about document shape and latency budget:
| Your workload | Right starting point |
|---|---|
| Direct lookups in short, clean documents | One-shot RAG |
| High-volume keyword or semantic retrieval | One-shot RAG |
| Answers live in known locations | One-shot RAG |
| Filings, contracts, manuals with tables and footnotes | Agentic Search |
| Scanned, table-heavy PDFs | Agentic Search |
| Answers spread across multiple documents needing reconciliation | Agentic Search |
| Interactive latency requirements (seconds) | One-shot RAG |
If a p90 of 154 seconds breaks your product, no accuracy gain fixes that. Batch analysis of financial statements is a different story.
Agentic Search Is Not the web_search Tool
Search for this product and Google surfaces Mistral's websearch documentation — a different thing. The web_search and web_search_premium tools belong to the Agents API and do live web retrieval with citations, billed at $30 and $50 per 1,000 calls respectively on Mistral's API pricing page, on top of model tokens. Agentic Search is the inverse direction: it digs into your own indexed corpus and does not query the open web. The two solve different problems, and only one of them has published pricing; no rates for Agentic Search appear anywhere in the announcement.
Two Ways to Start
- Search Toolkit. Open modules for ingestion, embeddings, and indexing, deployable in cloud or on-premises environments, with configurable parsers, chunking, embedding models, Vespa schemas, ranking profiles, and hybrid retrieval.
- Studio and Vibe Libraries. The managed route; a Search Starter App can build a local index over your corpus with default configuration, which is the fastest way to reproduce the benchmark setup before tuning anything.
Whichever route you take, run the same representative questions through one-shot RAG and the full loop, and compare answer quality, token use, and p90 latency before committing.
What's Still Unknown at Launch
Mistral has not published Agentic Search pricing, and the benchmark results are vendor-reported; no independent hands-on replication was available at launch. The earliest public reactions are first impressions:
"mistral just added agentic search to their api is the right move. wrapping perplexity-style search into a native agent tool saves us from writing fragile web-scraping pipelines and managing proxy rotation ourselves." — @AbdoKerdawy, AI product developer (X)
The token claims have already drawn pushback that Mistral's announcement does not fully answer:
"The Mistral agentic search retrieval tool claims to cut single-pass search failures by over 40%. If it bypasses database lookups, will it reduce the number of tokens you need to spend?" — @therawlogs, independent blogger (X)
The official answer is the 24–34% token reduction against the search-only loop. Whether that holds outside benchmark corpora is exactly what an independent run on the open FinanceBench set would settle.
Neighboring products face the same physics: Mixedbread's Toast-1 agentic search mode caps its loop at four rounds and eight parallel retrieval calls, and its own docs state plainly that it is slower than a single search. Latency is the tax the vendors in this category are paying for accuracy.
Direct Answers
Is Mistral Agentic Search available through the Agents API?
No. It ships through the Mistral Search Toolkit and Libraries in Studio and Vibe; the Agents API's web_search tool is a separate live-web product with separate pricing.
Does Agentic Search really reduce token usage?
Per the official benchmarks, the full navigation loop used 23.9–33.7% fewer tokens than a search-only agentic loop on FinanceBench; independent replication has not yet been published.
Which models work with Agentic Search?
Mistral tested Mistral Medium 3.5 and Z.ai GLM-5.2 with consistent gains, and describes the layer as model-agnostic with no fine-tuning required.
How much does Mistral Agentic Search cost?
No pricing has been published for Agentic Search itself. The $30 per 1,000 calls figure on Mistral's pricing page belongs to the older web_search agent tool.
The unresolved trade-off is simple: this is accuracy bought with minutes. For batch work over filings, contracts, and scanned government records, a +45 to +59 pp accuracy jump (OfficeQA Pro and FinanceBench full-loop results respectively) and a one-third token cut against a slower loop is a rational exchange. For anything user-facing, 71 seconds of mean latency is still disqualifying. And until someone outside Mistral reruns FinanceBench, the strongest numbers on the page remain the vendor's own.
Related reading: GLM-5.2 API guide and GLM-5.2 review cover the second benchmark model used in Mistral's tests.