AIREITER

OpenRouter Prompt Caching: Why Your Cache Isn't Hitting

Last Updated: 2026-08-22 01:24:54

OpenRouter's dashboard launch reported a platform-wide cache hit rate of 82.8% (@OpenRouter). Community threads tell a different story: sub-1% hit rates (@miolini) and bills 10–32x above expectations (r/openrouter). OpenRouter prompt caching does cut input costs, but only after four specific failure modes are fixed, and the biggest lever is keeping consecutive requests on one warm provider. One hard limit first: a prompt below the provider's minimum token floor will never cache, no matter what you configure.

What counts as a cache hit on OpenRouter

Prompt caching reuses a stable prompt prefix that a provider has already processed, so repeated input tokens get billed at a discount instead of full price. The cache lives on the specific provider endpoint that served the original request — which is why routing behavior matters as much as prompt structure. This is a different layer from response caching, which replays an identical complete request for free before routing ever happens.

Prompt cachingResponse caching
What it reusesStable prefix of any requestByte-identical request (SHA-256 of normalized body)
How to enableMostly automatic; cache_control for Anthropic, Qwen, GeminiX-OpenRouter-Cache: true header or preset
CostCached tokens at 0.1–0.5x inputHits free, misses billed normally
Lifetime3–5 min typical, up to 1h (Anthropic)Default 300s, range 1–86,400s
Blocks onPrefix change, provider switch, token minimumAny JSON change, API key rotation, account ZDR

Response caching shines for retries, unit tests, and repeated identical calls in agent workflows — JSON property order is part of the cache key, so a harmless serialization change is a miss. The canonical reference for the provider-side mechanics is OpenRouter's prompt caching guide:

OpenRouter prompt caching documentation page

What OpenRouter prompt caching costs, provider by provider

Cached reads cost a fraction of normal input price everywhere, but the write that creates the cache can carry a premium — 1.25x regular input on Anthropic for the default 5-minute TTL, 2x for the 1-hour option. Caching pays when the same prefix is re-read often enough to amortize the write; a one-off request can cost more with caching than without. On Claude Sonnet 4.6, cached input runs $0.30/M against $3.00/M fresh, per OpenRouter's own worked numbers.

Write and read multipliers by provider, from the same source:

ProviderCache writeCache readNotes
Anthropic1.25x (5 min) / 2x (1h)0.1xTTL selectable per breakpoint
OpenAI, pre-GPT-5.6Free0.25–0.5xAutomatic from 1,024 tokens
OpenAI GPT-5.6+1.25x0.25–0.5xExplicit breakpoints now supported
Google GeminiFree0.25xImplicit on 2.5+, ~3–5 min TTL
GrokFree0.25xAutomatic
MoonshotFree0.25xAutomatic
GroqFree0.5xKimi K2 models only
DeepSeek1.0x0.1xWrites billed as regular input
Alibaba Qwen1.25x0.1xExplicit cache_control required
Z.AIFree~0.2xCached storage listed as limited-time free

OpenRouter's tutorial models 10,000 repeated tokens over six turns: 6.0x a single turn uncached, 1.75x with Anthropic's 5-minute cache and sticky routing, 2.25x on a free-write provider at 0.25x reads. The model excludes growing messages and output tokens.

Relative input cost of 10,000 tokens over six turns under four caching setups

Anthropic's expensive write wins at six turns because its 0.1x reads dominate from turn two onward, and the gap widens as turns accumulate. The picture only flips when the 5-minute TTL expires between turns: you re-pay the 1.25x write on every request — 7.5x over six turns, worse than not caching — while a free-write provider at 1.0x input merely matches the uncached 6.0x.

Measure before you debug: three numbers that prove a hit

Every OpenRouter response carries the verdict in its usage object — cached_tokens, cache_write_tokens, and cache_discount, with field semantics documented in OpenRouter's caching guide. Reading these three before changing anything separates a real cache miss from a pricing surprise. cached_tokens above zero means the request hit a warm cache; zero means it did not, whatever the Activity dashboard suggests.

"usage": {
  "prompt_tokens": 10339,
  "prompt_tokens_details": {
    "cached_tokens": 10318,
    "cache_write_tokens": 0
  }
}

That response is a 99.8% hit — 10,318 of 10,339 prompt tokens came from cache. cache_write_tokens appears on the first cache-producing request; cache_discount reports the amount saved and can go negative on Anthropic writes, because the 1.25x write premium is a real cost that later reads pay back. You can pull the same numbers from the generation detail view in Activity (our activity dashboard walkthrough covers where) or from /api/v1/generation.

The raw metadata is the ground truth, not the UI. One SillyTavern user chased a phantom cache problem until they checked the logs directly:

"The raw openrouter metadata straight up says native_tokens_cached: 0 [and] usage_cache: null." — u/HauntingWeakness

If your three numbers say zero day after day, one of the four failure modes below is eating the cache.

The four ways a warm cache goes cold

OpenRouter's documentation and the community converge on four common causes of collapsed hit rates: sub-minimum prompts, TTL expiry between turns, a changed prefix, and provider drift. Each has a distinct signature in your logs and a different fix.

1. The prompt is under the provider's minimum

Providers that support prompt caching enforce a model-specific token floor — a 900-token system prompt will never cache on any Claude model, and padding it with filler is explicitly discouraged ("Don't pad the request with filler text just to force it," the OpenRouter tutorial says). The floors vary by a factor of four across the catalog:

Minimum cacheable prompt size by model family

Per OpenRouter's provider notes, Claude Opus 4.5–4.8 and Haiku 4.5 need 4,096 tokens before anything caches; Sonnet 4/4.5/4.6 and Opus 4/4.1 need 1,024; Gemini 2.5 Pro sits at 4,096 while Gemini 2.5 Flash takes 1,024; OpenAI models cache from 1,024. A short-prompt workload on Opus 4.8 is structurally uncachable — the fix is consolidating static material (tool schemas, reference docs, few-shots) into one prefix, or moving to a model with a lower floor.

2. The cache expired between turns

Anthropic's default cache lives 5 minutes; the 1-hour TTL costs a 2x write. Gemini's implicit cache survives roughly 3–5 minutes and — critically — reads do not reset the timer, per OpenRouter's tutorial. The sticky session that keeps you on the same provider dies after 10 minutes of inactivity. Agent loops that think for 5–6 minutes between calls blow through every one of these windows:

"OpenRouter [is] great for testing models. They're quietly terrible for production agents. The dirty secret? Caching is effectively zero in real workloads." — @ran_cohenn, describing 5–6 minute agent intervals expiring sticky affinity and landing full cache misses plus expensive cache writes

Anthropic's 1-hour TTL at 2x write beats re-paying 1.25x every five minutes, provided the session continues within the hour; with twenty-minute user gaps, no TTL on the menu survives, and caching only helps within a burst of turns.

3. The prefix changed on you

OpenRouter derives its default conversation key by hashing the first system message plus the first non-system message; anything that mutates the start of the prompt invalidates the cache from that point forward. The recurring offenders: RAG context injected above the system prompt, timestamps or request IDs embedded in the first message, tool definitions rewritten per call, and frontend chat apps inserting messages mid-history.

"Cache miss rate will increase if something at the beginning of the prompt is constantly changing." — u/Exact_Law_6489

Sometimes the change comes from tooling you didn't write. "I found Claude Code was causing cache hit problems for me, I think it's from how they inject tools," reports u/askchris. Gemini adds two traps of its own: OpenRouter uses only the last cache_control breakpoint you send, and the system instruction is treated as immutable cached content — dynamic material has to move into a later user message, not trail the system prompt. The fix in every case is the same discipline: static system prompt, tool schemas, and reference documents first; per-request variation last.

4. The request landed on a cold provider

OpenRouter routes across 70+ providers (per its own tutorial), and a prompt cache is local to the endpoint that wrote it. Sticky routing sends follow-ups back to the warm provider — but only when that provider's cache reads are cheaper than its regular input, and a manual provider.order overrides the stickiness entirely. A provider error also releases the pin.

The community data on this failure mode is stark:

  • @bruceforai measured the same model name across different providers and found cache hit rates from 95.3% down to 0%, with some third-party cache pricing 10x the official rate.
  • @Bryan_1269 got a very low hit rate on GLM 5.2 through OpenRouter and 85%+ on the identical prompt direct via Fireworks.
  • @miolini, on routing through OpenRouter: "cache hit rate is really bad, like less than 1%."

OpenRouter's official position is that pinning does hold — "when you get cached by a model or provider, you get pinned to it until the cache expires" (@OpenRouter) — which matches the docs and leaves provider variance, not pinning, as the thing to manage.

Where cache_control goes — and what strips it

Anthropic models on OpenRouter cache in two modes: one top-level cache_control object that advances automatically as the conversation grows (OpenRouter's recommendation for multi-turn chat), and explicit breakpoints on individual content blocks — up to four — for large fixed material like tool schemas, RAG documents, CSV dumps, or character cards. The top-level form works across Anthropic native, Vertex, Azure, and Bedrock, where OpenRouter translates it into a trailing breakpoint because Bedrock's API won't accept the top-level field. Setting an explicit TTL requires Chat Completions or the Anthropic Messages API, not Responses.

{
  "role": "system",
  "content": [
    {
      "type": "text",
      "text": "<20k tokens of tool schemas and reference docs>",
      "cache_control": { "type": "ephemeral", "ttl": "1h" }
    }
  ]
}

OpenAI works differently: caching is automatic from 1,024 tokens, and explicit prompt_cache_breakpoint markers exist only on GPT-5.6 and newer, set on an input_text or text block, with a 30-minute minimum TTL when you request one.

OpenRouter translates between dialects, per its provider notes: an Anthropic cache_control marker becomes an OpenAI breakpoint, an OpenAI breakpoint becomes a default 5-minute Anthropic marker, and TTL values never transfer. Qwen needs explicit cache_control markers, caches for 5 minutes, and supports them only on specific models (qwen3-max, qwen-plus, qwen3-coder-plus and others; snapshots like qwen3.5-plus-02-15 are excluded).

A quieter failure mode: some clients and gateways between your app and OpenRouter strip non-standard fields before forwarding:

"anthropic prompt caching dropping to zero behind gateways is usually a marshalling bug. ... cache_control markers being silently stripped before forwarding to openrouter. you cannot abstract providers by dropping their schema extensions." — @SiddharthInk_

Verify the marker arrives: inspect the raw request metadata on the Activity generation detail, or send one test request with curl where nothing in between can interfere. A tool that flattens messages into a single blob destroys breakpoints no matter how correctly you placed them. OpenRouter's examples repo has runnable TypeScript, Vercel AI SDK, and Effect samples that keep the markers intact.

Pin the provider: session_id and provider order

A stable session identity is the strongest routing lever: session_id pins follow-up requests to the provider that served the first successful request, before any cache hit is observed. Without it, stickiness begins only after the first detected cache hit, and the default identity — a hash of the first system message plus first non-system message — silently rewires on any prefix mutation (failure mode 3), per OpenRouter's routing documentation.

{
  "model": "anthropic/claude-sonnet-4.6",
  "session_id": "user-8801-thread-3",
  "messages": [ ... ]
}

Mechanics worth knowing: session_id goes in the request body or the x-session-id header (body wins if both are set; 256-character max; OpenRouter falls back to the OpenAI-style prompt_cache_key if neither is present).

Two caveats from the docs: provider errors release the pin, and Batch API lines execute concurrently and out of order, so a cache write from one line isn't visible to the next — share a "ttl": "1h" prefix across batches or warm it with one synchronous request first. (The auto router guide covers Auto Router's best-effort reuse of the resolved model.)

If pinning alone isn't enough, restrict the provider set outright:

"The solution that I've found is setting a preferred list of providers to be used in order of preference." — u/nabil9506

A provider.order list of two or three providers with cheap cache reads trades failover breadth for cache locality — a reasonable trade for agent workloads. u/welcome_to_milliways calls the manual setup burden "a pretty fundamental flaw in OR"; fair or not, it's the current contract.

When caching through a router isn't worth it

Prompt caching through OpenRouter stops paying in three recognizable situations: prompts that never reach the model's token floor, sessions with gaps longer than every available TTL, and one-off requests whose write premium is never amortized by a discounted read. There's also a fourth: tooling you can't change that strips cache_control before it reaches the router. @grapeot puts the stakes in perspective — when caching fails at the gateway layer, the cost gap is an order of magnitude, dwarfing the routing fee itself.

For cache-critical workloads where none of the fixes apply, a single fixed upstream beats a router: deterministic cache behavior, no pinning to manage. A direct Claude API endpoint with Anthropic's own caching is the straightforward escape hatch when provider drift is unfixable.

Account-level Zero Data Retention disables response caching entirely; for prompt caching under ZDR, OpenRouter's analysis of whether implicit caching counts as data retention is the reference to check.

The fix order

Debugging in measurement order recovers most of the savings with the least churn — verify first, then work down the stack from prompt to routing to TTL:

#ActionWhat it settles
1Read cached_tokens and cache_discount on a few real requestsHit-rate problem vs. pricing expectation problem
2Compare prompt size against the model's token floorRules out "never cacheable" before anything else
3Freeze the prefix: static system prompt, schemas, docs first; timestamps and RAG lastKills the silent invalidation class
4Pass session_id on every request in a conversationProvider pinning from turn one, not after first hit
5Set provider.order to two or three cheap-cache-read providersRemoves cross-provider drift
6Add "ttl": "1h" (Anthropic) or switch to a free-write provider for long sessionsHandles expiry between turns

Steps 1–3 clear the failure classes you control in code; steps 4–6 reconcile the sub-1% reports with the 82.8% headline. Related reading: the OpenRouter pricing guide for how cached tokens land on your bill, the auto router guide for model-pinning behavior, and the activity dashboard guide for monitoring hit rates over time.