A half-price API is attractive until the result arrives after the deadline. OpenRouter’s Batch API is a good fit for offline text and embedding jobs, but not for interactive calls: it is asynchronous, has a 24-hour completion window, and its headline discount does not apply identically to every charge.
The one-line buying decision
Use OpenRouter Batch API for labeling, evaluations, embeddings, backlog summarization, and other durable jobs that can wait. Keep user-facing chat, IDE agents, web-search workflows, and multimodal requests on the synchronous API.
OpenRouter says Batch generally offers about 50% lower per-token pricing across more than 70 models. The formal completion window is 24 hours. In its launch announcement, OpenRouter reported a 7-minute median and 90% completion within one hour during beta; those figures are observations, not an SLA (official announcement).
What the 50% discount actually covers
The discount applies primarily to model token pricing. It is not a blanket reduction for every component of an inference bill.
| Cost or control | Batch API treatment |
|---|---|
| Input and output tokens | Typically about 50% of standard model pricing |
| Web-search calls | Billed at standard rates, according to the official quickstart |
| Prompt caching | Varies by model; check the model page |
| BYOK inference | Provider bills inference directly; OpenRouter reports its BYOK fee separately |
| Exact eligible price | Confirm on the individual model page and completed batch usage |
Will Cygan’s batching cost walkthrough shows a Claude Sonnet 5 example dropping from $40 synchronous to $20 batch for 10 million input and 2 million output tokens. That is a model-specific calculation, not a universal quote.
“The batch route bills at exactly half the sync rate.” — Will Cygan, Batching (LLM Inference)
Do not budget the saving before checking the provider and model. A real user, @fogelmania, reported that one beta model was more expensive than concurrent synchronous calls because its batch traffic reached a different provider: @fogelmania’s post. The report is a warning to inspect completed cost, not proof that every model behaves that way.
A batch is a job, not a faster endpoint
OpenRouter’s Batch API announcement and quickstart describe a job workflow rather than an immediate completion. A successful submission returns HTTP 202 Accepted and a batch ID with status validating. The normal lifecycle is:
validating → in_progress → finalizing → completed
The other terminal outcomes are failed, expired, and cancelled. Your worker should persist the batch ID and poll until a terminal state instead of holding an interactive request open.
OpenRouter reported more than 230,000 beta batches, with a 7-minute median and 90% completing within an hour. A user test by @luismmolina reported 5–8 minutes at one point on launch day (test post); these observations do not replace the 24-hour planning boundary.
The implementation shape that avoids rework
The current quickstart uses an inline JSON requests array rather than a JSONL file upload. Each row needs a unique custom_id; that ID matches a completed answer or error to the original record.
A minimal request shape is:
{
"endpoint": "/v1/chat/completions",
"model": "openai/gpt-4o",
"requests": [
{
"custom_id": "ticket-0001",
"body": {
"messages": [
{"role": "user", "content": "Classify this ticket: ..."}
]
}
}
]
}
The quickstart documents POST https://openrouter.ai/api/beta/batches. The top-level endpoint and model apply to the whole batch, so different API shapes or models require separate batches. Supported shapes include Chat Completions, Responses, Anthropic Messages, and Embeddings.
After submission, poll GET https://openrouter.ai/api/beta/batches/:id. A completed batch returns results inline. Each result contains either a response or an error, and request_counts separates total, completed, and failed rows. Retry failed rows by their custom_id; do not automatically replay the entire batch.
If provider behavior matters for data policy, BYOK, or URL assets, pin the provider using the documented provider controls rather than relying on cheapest-provider routing. Verify that the selected model and provider expose an eligible batch route before deployment.
Where Batch breaks the workflow
The quickstart’s limitations make Batch a text-first workflow. It rejects images, audio, video, and file content parts in batch requests. Base64 and data: URI assets are rejected; supported URL assets depend on the provider. OpenRouter’s own web-search plugin is unavailable in Batch.
Use the synchronous API when a user is waiting, a model must inspect a local upload, a request needs audio or video, or the application requires a second-level response-time target.
Cost example: when the saving is real
Consider 10,000 support tickets, each using 1,000 input tokens and 200 output tokens. That is 10 million input tokens and 2 million output tokens.
| Route | Input | Output | Total |
|---|---|---|---|
| Synchronous example | 10M × $2 = $20 | 2M × $10 = $20 | $40 |
| Batch example | 10M × $1 = $10 | 2M × $5 = $10 | $20 |
The nominal saving is $20 per run, or $1,040 per year if this example runs weekly. The effective saving is lower whenever recovery, monitoring, or an emergency synchronous fallback costs more than the nominal difference.
Build that reserve into the decision. If a deadline is hard, compare the 24-hour window with the remaining time for a reduced-scope rerun or synchronous fallback. A batch that is cheaper per token but unusable after the deadline is not cheaper for that business process.
FAQ
Is OpenRouter Batch API always half price?
No. OpenRouter describes the discount as typical and model-dependent. Web-search charges remain standard, caching varies, and BYOK separates provider inference costs from OpenRouter fees.
How long does an OpenRouter batch take?
The supported completion window is 24 hours. Beta timing reports are useful context, not a guaranteed service level.
Can I upload JSONL or mix models?
The quickstart accepts an inline JSON requests array. The model and API shape apply to the whole batch, so separate batches are required for different models or endpoint formats.
Can I retry only failed rows?
Yes, when a completed batch returns row-level errors, use each row’s custom_id to construct a smaller retry batch. Treat a batch-level failure, expiration, or cancellation separately because results may be unavailable.
Should I use Batch or synchronous API?
Choose Batch for non-urgent background work. Choose synchronous inference when the result is part of an active user interaction or requires unsupported modalities and tools.
The practical call: use Batch selectively
Before moving a workload, check five things:
- The model page shows an eligible batch route and the expected provider.
- The business process can tolerate the full 24-hour window.
- Every row has a stable
custom_idand a retry plan. - The application records actual completed usage and cost.
- Inputs and results have an owner and cleanup policy.
OpenRouter’s quickstart says batch inputs and results are retained for 30 days unless deleted earlier. Delete terminal batches when the artifacts are no longer needed.
The best first migration is a frozen, reviewable corpus, not a customer-facing path where a late answer costs more than the token discount saves.