The same JSON schema that returns clean, typed output from one OpenRouter model can come back with different keys, an empty string, or a 400 error from the next one, with the same request body. Reddit user u/MicBeckie tested Qwen models through OpenRouter's structured outputs and the result was "9 times out of 10 I always got errors," while OpenAI models in the same setup held the schema.
That gap is not a bug you can file. OpenRouter's structured output support is determined per endpoint, not per model, and "support" spans three enforcement tiers, from native strict schema enforcement down to providers that treat your schema as a suggestion. This guide covers how the feature routes, the six ways requests fail in practice, and the hardening steps that make schema output shippable. Enforcement mechanics follow the official structured outputs docs; the failure evidence comes from developer threads linked inline.
What "structured output support" means on OpenRouter
OpenRouter accepts a response_format parameter with type: "json_schema", a schema name, a strict flag, and the JSON Schema itself. A minimal request looks like this:
{
"model": "openai/gpt-4o",
"messages": [{ "role": "user", "content": "Extract the shipping info" }],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "shipping_info",
"strict": true,
"schema": {
"type": "object",
"properties": {
"tracking_number": { "type": "string", "description": "Carrier tracking ID" },
"carrier": { "type": "string" },
"eta_days": { "type": "number", "description": "Days until delivery" }
},
"required": ["tracking_number", "carrier", "eta_days"],
"additionalProperties": false
}
}
}
}
Two details from the official docs decide whether this works at all:
- Support is per endpoint, not per model. A model served by five providers may have structured outputs working on two of them. The model page's Providers section exposes a
structured_outputsparameter per provider, and the docs warn that "endpoint support can also change over time." - Coverage grew from a narrow base. OpenRouter announced structured outputs on December 12, 2024 with support for OpenAI 4o and Fireworks models only. Everything else arrived later, provider by provider, so any model list written down today ages fast.
The docs also recommend attaching descriptions to every property and setting additionalProperties: false, because on the lower enforcement tiers the schema doubles as prompt material.
The three enforcement levels behind one flag
strict: true means different things depending on where the request lands. The official guide splits provider behavior into three tiers:
| Tier | What the provider does with your schema | Can you trust the output? |
|---|---|---|
| Native strict mode | Enforces the schema exactly at decode time | Yes: output matches schema by construction |
| Translated format | Converts your schema into a provider-specific structured-output format | Mostly: limited to the schema features that format supports |
| Strong hint | Injects the schema as guidance for the model | No: schema-shaped on a good day, hallucinated keys on a bad one |
OpenRouter does not label which tier a given endpoint uses at request time; the docs defer to each provider's own documentation. Native strict modes also restrict which JSON Schema features they accept, so exotic keywords can fail on the strictest endpoints while passing as hints elsewhere.
One Claude special case is documented on the provider routing page: for response_format.type: "json_schema", OpenRouter automatically applies Anthropic's structured-outputs-2025-11-13 beta header, which enables strict, schema-validating tool arguments. But for strict: true tool definitions sent as tools, the caller must send that beta header explicitly. Otherwise OpenRouter strips strict and routes the request without it. The failure mode is silent: your tool calls stop being schema-validated, and nothing errors.
The six ways the same schema fails
Two failure classes fail fast, documented in the official guide; four more show up in community threads, and they are the ones that eat afternoons.
Fail-fast class 1: the endpoint doesn't support structured outputs. The request errors out stating the capability is unsupported. Annoying, but unambiguous. Fail-fast class 2: your JSON Schema is invalid. The API rejects the request because the schema itself doesn't parse or violates the endpoint's schema rules.
Silent failure 1: the schema gets ignored. The response is valid JSON for a different schema. From the schema-not-followed thread on r/LocalLLaMA, u/DaniyarQQQ:
It returns json that does not look like my schema at all.
u/MicBeckie, on the diagnosis problem in the same thread:
I either see successes when the json corresponds exactly to the requirement, or I get an error without being able to look at the json.
Wrapper failure 2: a 400 about tool_choice you never sent. In the linked LangChainJS case, withStructuredOutput() implemented "structured output" by forcing tool_choice to a generated function. On models that advertise tool calls but don't support forced tool choice, the request dies with invalid_request_error; in the DeepSeek v4 case the error named the model directly — deepseek-reasoner does not support this tool_choice. u/shansoft hit exactly this via LangChainJS (thread), where u/eyueldk's assumption "it says it supports tool calls, thus should support structured output" turned out not to hold. Tool-call support and strict schema support are separate capabilities.
Silent failure 3: no error, no content. A gpt-oss-120b report describes a strict-schema request that 400s on a direct provider route but returns, through OpenRouter, a 200 with an empty message.content. Another thread on r/openrouter shows a "supported" model returning only [1] or [1.1]. An SDK that happily parses an empty string pushes the failure three layers downstream.
Silent failure 4: the endpoint hangs. u/Beneficial-Loss-1031, on endpoints that did advertise structured output for DeepSeek v4 (thread):
deepinfra/fp4andakashml/fp8have structured output option, but I waited for 3 min on each for the API to return, but didn't get anything.
| # | Failure form | What you see | Typical cause |
|---|---|---|---|
| 1 | Unsupported endpoint | Error: structured outputs not supported | Routed to a provider without the capability |
| 2 | Invalid schema | API error on the request | Schema breaks the endpoint's rules |
| 3 | Schema ignored | Valid JSON, wrong keys | Hint-tier enforcement |
| 4 | tool_choice 400 | invalid_request_error | SDK emulating schema via forced tool call |
| 5 | Empty content | 200, empty message.content | Provider mishandling strict mode |
| 6 | Hang | No response for minutes | Unconfirmed in the report — a 3-minute wait on fp4/fp8 endpoints |
Harden the request before you blame the model
The single highest-leverage setting is require_parameters: true in the provider object. By default it is false, and unknown parameters are passed through to providers that silently ignore them. Even at false, response_format and structured outputs act as a soft preference among endpoints: preferred, not guaranteed. Setting the flag to true restricts routing to endpoints that support every parameter you sent, per the provider routing docs:
{
"model": "deepseek/deepseek-chat",
"messages": [{ "role": "user", "content": "Extract the shipping info" }],
"response_format": { "type": "json_schema", "json_schema": { "name": "shipping_info", "strict": true, "schema": { "...": "..." } } },
"provider": {
"require_parameters": true,
"order": ["fireworks"],
"allow_fallbacks": false
}
}
Every restriction shrinks the eligible provider pool, and allow_fallbacks: false trades availability for determinism. The same routing docs describe the default strategy as load-balancing by uptime and the inverse square of price over the preceding 30 seconds, which optimizes for cheap and healthy rather than schema-capable. Pinning order to one provider and disabling fallbacks makes routing reproducible: the request can no longer drift to a different provider mid-outage. The enforcement tier of that one endpoint is still yours to verify.
Two audit habits catch what routing can't:
- Check which provider served the request. OpenRouter's generation metadata exposes provider routing per generation, alongside model, latency, and token counts. If output quality drifts, that attribution tells you whether the model changed behavior or the router changed providers.
- Validate client-side regardless. None of the tiers substitute for a Pydantic or Zod parse on your side. The recurring lesson from the r/LLMDevs testing threads is that "valid JSON," "schema-valid," and "semantically correct" are three different bars, and only the first two are even partially the API's job.
Streaming works, but the parsing is on you
Structured outputs compose with stream: true. The docs describe the contract as the model streaming valid partial JSON, with the assembled response matching the schema once the stream completes. That conformance inherits the endpoint's enforcement tier — a hint-tier endpoint can still assemble non-conforming output — so validate the final object yourself. The docs don't ship an incremental parser either; for latency-sensitive UIs that gap is the real engineering problem. From the streaming best-practice thread on r/LLMDevs:
I just ended up writing a function that completes the JSON myself. — u/am174744
"…it's an actual state machine." — u/ImNotLegitLol, correcting the repair-then-parse mental model
Practical options: parse partial JSON with a streaming-tolerant parser, render only completed fields, or skip incremental rendering and show a spinner until the final object assembles.
What Response Healing will and won't fix
OpenRouter's Response Healing plugin targets non-streaming json_schema requests and repairs imperfect formatting: truncated JSON, stray markdown fences, that class of problem. Two limits matter more than what it fixes:
- Streaming is out of scope. The docs scope the plugin to non-streaming requests.
- Schema violations are out of scope. Healing makes JSON parseable; it does not make a response that ignored your schema conform to it. Failure form 3 above is untouched.
Picking models that actually honor schemas
Model lists rot; screening criteria don't. Three filters catch most of the failure forms above:
- Native strict enforcement. Prefer models whose serving provider enforces schemas at decode time over providers that translate or hint. The model page's Providers table shows which endpoints advertise
structured_outputs; the provider tier determines enforcement quality. - One auditable provider. Cross-check provider attribution against a known-good endpoint across multiple calls. If the router spreads requests across providers on different tiers, your failure rate is a routing lottery. Pin the provider or pick a single-provider model.
- A smoke test you ran, not one you read about. Community signals age fast in both directions: the Qwen error reports above and DeepSeek v4's missing support may both change as providers update endpoints. The only reliability number that matters is the one your own schema produces.
OpenRouter structured outputs FAQ
What's the difference between json_object and json_schema?
json_object only asks for syntactically valid JSON; json_schema supplies a schema the response must conform to. json_object guarantees JSON syntax, not conformance to your field-level schema, so validate it yourself when downstream code needs named fields.
Which OpenRouter models support structured outputs?
There is no static list to trust: support is per endpoint, it changes over time, and it started with OpenAI 4o and Fireworks models only in December 2024. Check the model page's Providers section for the structured_outputs flag on each endpoint.
Why does the model ignore my schema?
Three common causes: the request routed to a hint-tier or non-supporting endpoint (fix with require_parameters: true and provider pinning); the schema relies on keywords the endpoint's strict mode rejects; or an SDK wrapper is emulating structured output through tool calling on a model that doesn't support forced tool choice.
Can I use Pydantic or LangChain with OpenRouter structured outputs?
Yes. The official docs describe the request format as compatible with OpenRouter's chat-completions-style API, so Pydantic-generated schemas and the OpenAI SDK work directly. LangChain's withStructuredOutput() also works, but verify it's sending response_format rather than emulating via tool_choice, which is what produced 400 errors on DeepSeek v4.
Does structured output work with streaming?
Yes. The stream emits valid partial JSON, but final schema conformance depends on the endpoint's enforcement tier — validate the assembled object yourself. Incremental parsing of fragments is your application's job, and Response Healing does not apply to streams.
Does OpenRouter validate responses against my schema?
Not as a guarantee across all endpoints: enforcement depends on the provider tier, and Response Healing only repairs malformed JSON, not schema violations. Client-side validation remains mandatory.
The 10-call smoke test
Before any model goes into production behind structured outputs, run this:
- Fix one representative schema: medium complexity,
additionalProperties: false, descriptions on all properties. - Send 10 identical requests with
strict: trueandrequire_parameters: true, fallbacks enabled — this pass deliberately tests fallback behavior, so leave them on. - Score each response on three bars: parseable JSON? Schema-valid? Semantically sane?
- Record which provider served each response, via the generation metadata. A 10/10 pass rate served by four different providers is a routing lottery, not a guarantee.
- Decide: ship as-is, pin
provider.orderto the endpoint that passed and re-run the 10 calls pinned, or swap models and add a client-side validation-and-retry layer.
The pass threshold is yours to set, but anything below 9/10 on a fixed schema means retries and validation code are not optional. They're the product.
Related reading: how OpenRouter's auto router picks providers, cutting costs with OpenRouter prompt caching, and fixing OpenRouter 429 rate limits.