The best OpenRouter roleplay model depends on whether the priority is predictable voice, a long context, or a zero-dollar session. In the September 2, 2026 snapshot, DeepSeek V4 Flash 0731 is the paid cost-and-context baseline to test first: OpenRouter lists a 1M-class context and $0.05 input / $0.16 output per million tokens. Free routes are useful for testing, but their availability and capacity can change.
This guide focuses on technical behavior and cost for non-explicit, fictional roleplay. It does not claim a universal “uncensored” winner or publish refusal percentages without a controlled, apples-to-apples run.
Which OpenRouter roleplay model should you pick in 2026?
Pin one paid model and one live free fallback; treat the usage leaderboard as a shortlist signal, not a quality score.
| Use case | First model to check | Why it fits | Main limitation |
|---|---|---|---|
| Paid cost-and-context baseline | DeepSeek V4 Flash 0731 | 1M-class context, $0.05/$0.16 per 1M tokens, and current roleplay traffic on OpenRouter | Usage is not controlled RP testing |
| Free-first trial | MiniMax M3 (free) | Listed in OpenRouter’s live free collection with a roughly 1M context | Free capacity, limits, and availability can change |
| Low-cost paid comparison | GLM 5.3 Flash | 1M-class context, 22 listed providers, and a temporary $0.075/$0.25 price | The discount is time-limited; the model is positioned for agentic work |
| Mid-price comparison | MiMo-V2.5 in OpenRouter’s roleplay collection | $0.119/$0.238 per 1M tokens and 1.05M context in the live listing | The cited page provides catalog and usage data, not a controlled RP score |
| Automatic free fallback | openrouter/free | Selects an available free route without a fixed slug | The underlying model can change between requests |
Start with DeepSeek as the paid baseline, confirm MiniMax M3’s current free endpoint before relying on it, and use GLM 5.3 Flash as a second paid control. The shortlist is based on current usage, price, context, and routing data, not a universal prose ranking.
What OpenRouter’s roleplay ranking actually tells you
OpenRouter’s Roleplay and Creative Writing collection is a live usage-based discovery page. Its September 2, 2026 capture places DeepSeek V4 Flash first.
The ranking measures usage, not character-card adherence, name retention, prose quality, or refusal rate. A popular model may simply be cheap, fast, widely available, or easy to route.
Build a shortlist from the collection, then run the same safe card and opening scene through two pinned model IDs. That test answers a narrower question: which route works better for your prompt, settings, and budget?
The current shortlist: context, price, and availability
The numbers below come from OpenRouter’s roleplay, free-model, and model pages checked on September 2, 2026. Prices and free status are operational snapshots.
| Model ID | Context shown | Input / output price per 1M | Free route seen? | What the data supports |
|---|---|---|---|---|
deepseek/deepseek-v4-flash-0731 | 1,310,720 | $0.05 / $0.16 | Check the live catalog | Low listed paid price, 30 providers, and automatic failover |
minimax/minimax-m3 | 1,048,576 | $0.23 / $0.96 paid | Yes, listed as MiniMax M3 (free) | Large context and zero-cost testing when the free endpoint is available |
z-ai/glm-5.3-flash | 1,310,720 | $0.075 / $0.25 promotional | Check the live catalog | 22 providers and 99.89% three-day availability in the snapshot |
xiaomi/mimo-v2.5 | 1.05M | $0.119 / $0.238 | Not established from the cited page | Roleplay-collection usage signal and moderate paid pricing |
openrouter/free | Router page: 200K | $0 | Router, not a fixed model | Convenient for experiments; random selection reduces reproducibility |
DeepSeek V4 Flash 0731: the cost-and-context baseline
DeepSeek V4 Flash 0731 has a 1,310,720-token context, up to 393,216 completion tokens, and listed pricing of $0.05 input / $0.16 output per million tokens. OpenRouter’s page lists 30 providers and describes Balanced, Nitro, and Exacto routing modes.
Use the same-card check before committing a long campaign because the model page does not publish a measured character-consistency score.
MiniMax M3: the free trial candidate
OpenRouter’s free-model collection lists MiniMax M3 at $0 for input and output, with a 1.05M context display. The paid MiniMax M3 page lists 1,048,576 context, text/image/video input, text output, and paid pricing of $0.23 input / $0.96 output per million tokens.
OpenRouter’s Free Variant documentation says a :free route may have different availability and rate limits from the paid model. Select the exact free slug shown in the live catalog when reproducibility matters; use openrouter/free only when automatic selection is acceptable.
GLM 5.3 Flash: the provider-diversity option
GLM 5.3 Flash has a 1,310,720-token context window and is listed at $0.075 input / $0.25 output per million tokens through a 50% promotional discount via Z.ai through September 9, 2026 at 16:00 UTC. OpenRouter’s page lists 22 providers and 99.89% three-day availability.
The model page positions GLM 5.3 Flash for efficient coding and long-horizon agent tasks, not specifically for roleplay. Test its RP fit with the same prompt rather than inferring it from the product description.
MiMo-V2.5 and other fallbacks
MiMo-V2.5 appears in OpenRouter’s roleplay collection with a 1.05M context and $0.119 input / $0.238 output pricing. It is a reasonable mid-price control, but the cited page does not publish a controlled roleplay result. Save a second exact slug before a session fails; a preselected fallback is more predictable than a random model swap.
The 1,000-round cost test: context management changes the bill
A roleplay round is one user message plus one assistant reply. OpenRouter bills tokens, so there is no universal price per 1,000 rounds; the amount of history resent by the frontend determines much of the bill.
This planning model assumes 1,000 rounds, 200 new user tokens per round, 500 assistant output tokens per round, and a 2,000-token fixed character/system prompt. In the full-history case, the 1,000 requests resend the growing conversation, producing about 351.85M input tokens and 0.5M output tokens in total.
| Model | Compact-history scenario: 4M input + 0.5M output | Full-history scenario: 351.85M input + 0.5M output |
|---|---|---|
| DeepSeek V4 Flash 0731 ($0.05/$0.16) | $0.28 | $17.67 |
| GLM 5.3 Flash ($0.075/$0.25) | $0.43 | $26.51 |
| MiMo-V2.5 ($0.119/$0.238) | $0.60 | $41.99 |
| MiniMax M3 paid ($0.23/$0.96) | $1.40 | $81.41 |
| MiniMax M3 free | $0.00, subject to limits | $0.00, subject to limits |
The compact-history column averages 4,000 input tokens per request. Full-history resends the same early messages repeatedly, which is why a large context window does not make a 1,000-round campaign cheap.
Under these assumptions, the final full-history prompt is about 701,500 input tokens, before extra formatting and metadata. A 1M context window can fit this example, but longer replies, larger cards, lorebooks, or hidden templates can exceed it. Context capacity is not automatic memory.
Prompt caching may reduce repeated-input cost when the provider and request path support it. Check OpenRouter’s Activity page instead of assuming every frontend receives the same cache treatment. OpenRouter’s pricing guide explains the broader token-billing model.
What a safe consistency check can—and cannot—tell you
A non-explicit comparison can use a fictional adult character, a benign scene, and identical settings for every model:
- Put three durable facts in the character card, such as a profession, location, and non-sensitive goal.
- Require third-person narration or one clearly defined point of view.
- Run 20 turns with the same opening and a 400–600-token output cap.
- At turns 5, 10, and 20, ask the character to act on one earlier fact without restating it.
- Record name changes, contradictory facts, unwanted point-of-view switching, repetition, empty responses, and refusals.
This compares your prompt and route, not universal model quality. A refusal can come from the model, provider, frontend classifier, or prompt, so this guide reports no refusal rate. For speed, record first-token latency and generated tokens per second across three identical runs, noting the model and provider route.
How long can OpenRouter’s free tier support roleplay?
OpenRouter’s official SillyTavern guide states that accounts with less than $10 in purchased credits receive 50 free-model requests per day. After at least $10 in purchased credits, the guide states the allowance becomes 1,000 requests per day. It also lists a 20-request-per-minute limit.
If one assistant generation equals one round, that is approximately 50 first-pass rounds or 1,000 first-pass rounds per day. Regenerations, retries, failed requests, and multi-message workflows consume requests too, so completed scenes may be fewer.
The two free options behave differently:
model-name:freepins a named model’s free variant. It is more reproducible, but the endpoint can disappear or be rate-limited.openrouter/freeselects an eligible free model automatically. OpenRouter’s Free Models Router documentation says the selection is random after capability filtering, and the response identifies the underlying model. It is convenient for experiments, but the character voice can change between requests.
Check the live OpenRouter free-model collection before configuring a frontend. Do not build a long campaign around a free slug copied from an old guide.
SillyTavern setup for a stable OpenRouter session
The official OpenRouter workflow uses Chat Completion rather than Text Completion. A successful key connection still does not guarantee that the selected model, provider route, or context size can generate a response.
- Open SillyTavern’s API Connections panel.
- Select API Type: Chat Completion.
- Select Source: OpenRouter.
- Authorize through the available flow or paste an API key created in OpenRouter’s key settings.
- Click Connect, select the exact model slug, and send a Test Message.
- Save one connection profile for the paid model and another for the free fallback.
Start with a context size supported by the exact endpoint, streaming enabled, and a moderate output cap. Keep the character card compact and summarize old turns before the prompt approaches the endpoint limit. Test sampling settings separately because they can change output length and pacing.
If the dropdown is empty, reconnect and refresh the model list. If connection succeeds but generation fails, check the slug, credits, provider availability, and context overflow before rewriting the prompt.
JanitorAI setup and error fixes
OpenRouter’s API is OpenAI-compatible. The standard chat-completions endpoint documented in the OpenRouter Quickstart is:
https://openrouter.ai/api/v1/chat/completions
The article does not rely on a current JanitorAI UI capture, so use the following as a generic OpenAI-compatible integration pattern and verify the fields shown in your current JanitorAI screen:
- Create an OpenRouter API key and keep it private.
- Choose the current OpenAI-compatible or custom API option, if available.
- Enter the full chat-completions URL when the field expects a completion endpoint.
- Paste the exact current OpenRouter model ID, including
:freefor a free variant. - Save, refresh JanitorAI, and test with a short non-explicit message.
| Symptom | Likely cause | First fix |
|---|---|---|
| 401 or invalid token | Key is wrong, expired, or pasted with extra characters | Regenerate the key and paste it again; do not use a shared public key |
| 404 or no endpoint | Model slug or URL is outdated | Copy the current slug from OpenRouter and verify the full endpoint |
| 429 | Free-model daily/per-minute limit or provider overload | Wait, reduce retries, switch to the pinned fallback, or use a paid route |
| Empty response | Unsupported route, output limit, or context overflow | Lower context/output settings and test the model directly in OpenRouter |
| The bot forgets the opening | Frontend truncated old history | Reduce prompt size, summarize older turns, and match the endpoint context |
| Refusal or sudden style change | Provider policy or a different fallback route | Check the actual model/provider and use a neutral test prompt; “uncensored” is not a guarantee |
Privacy, provider routing, and compliance
OpenRouter’s FAQ says it records basic request metadata such as timestamps, model, and token counts, while prompts and completions are not logged by default unless the user opts in. Upstream providers can have their own retention and training policies.
Zero Data Retention can restrict routing to providers that meet the selected retention policy, but it may remove some free routes. Free access is not a privacy guarantee, and “uncensored” is not a stable technical property: providers can apply their own policies, and routing can change which provider serves a request.
Use non-explicit fictional prompts and do not attempt to bypass provider safeguards.
FAQ
What is the best OpenRouter model for roleplay right now?
DeepSeek V4 Flash 0731 is the paid cost-and-context baseline to compare in this snapshot. It combines current OpenRouter roleplay usage evidence with a 1M-class context and $0.05/$0.16 per-million-token pricing, but it is not a controlled universal quality score.
What is the best free OpenRouter model for roleplay?
There is no permanent winner: MiniMax M3 is the free-first candidate in the live collection checked here, while openrouter/free is more convenient but less reproducible.
How much does 1,000 rounds of roleplay cost?
With the assumptions in this guide, compact-history usage costs about $0.28 on DeepSeek V4 Flash 0731 and $0.43 on GLM 5.3 Flash; full-history usage is modeled at $17.67 and $26.51 respectively.
Does a larger context window prevent the model from forgetting?
No. Context is the maximum the endpoint can accept, not a guarantee that every detail will be used correctly; summaries and recent-turn limits still matter.
Why did my roleplay model refuse or change behavior?
The provider, fallback route, free endpoint, or practical context limit may have changed, so verify the actual model ID and route before treating one refusal as a model-wide property.
Run the same safe 20-turn card through DeepSeek V4 Flash 0731 and one current free endpoint, then save both as named profiles. The paid route offers more predictable pricing and availability; the free route offers experimentation at the cost of changing capacity and model identity.