The config file that ships with Baidu's model caps max_position_embeddings at 32,768. That number is where the name stops being literal, and it is the most useful thing to know before you plan a deployment: Unlimited OCR turns a page-by-page OCR loop into one forward pass, but only for documents that fit a 32K token budget and only when the scans are reasonably clean.
The rest of the release is stronger than that caveat suggests. The weights are MIT licensed, the safetensors metadata reports 3,336,106,240 parameters in BF16, and the model scores 93.23 overall on OmniDocBench v1.5 against 87.01 for the DeepSeek-OCR baseline it was built from. Downloads on Hugging Face stood at 2,694,935 for the trailing 30 days on 2026-07-29, against 3,389 likes.
Where the name stops being literal
In the multi-page path, each page is encoded at 1024×1024 and compressed 16 times down to roughly 256 visual tokens. That gives you a page budget you can compute before writing any code.
| Pages in one pass | Visual tokens (prefill) | Tokens left for output | Budget per page |
|---|---|---|---|
| 10 | ~2,560 | ~30,200 | ~3,020 |
| 20 | ~5,120 | ~27,600 | ~1,380 |
| 40 | ~10,240 | ~22,500 | ~560 |
A slide deck or a sparse contract fits 40 pages comfortably. A dense two-column newspaper page can run past 560 markdown tokens on its own, so the practical ceiling drops well below 40 pages for that kind of input. The paper says this outright: parsing cannot be truly unlimited under a finite context length, because prefill still grows with page count. Baidu's stated roadmap is a 128K context version plus a "prefill pool" that fetches page chunks on demand. What the name describes is unlimited decode length relative to cache size, not unlimited pages.
What R-SWA changes, and what it leaves alone
Two things changed from DeepSeek-OCR. The DeepEncoder vision stack (a SAM-ViT-B plus CLIP-L cascade) was kept and frozen during training. Every decoder attention layer was swapped for Reference Sliding Window Attention, where each generated token attends to all reference tokens (the visual tokens plus the prompt) and to only the last 128 output tokens. The config.json confirms it: sliding_window_size: 128, 12 decoder layers, 64 routed experts with 6 active per token.
The gains are broad rather than narrow, and they land on the parts of a page that survive longer generation: on OmniDocBench v1.5, formula CDM went from 83.37 to 92.61 and table TEDS from 84.97 to 90.93, with reading-order edit distance halving from 0.086 to 0.045. Since the encoder was not retrained, those deltas come from the decoder change rather than from a stronger vision stack.
It also means anything the encoder was bad at, the new model is still bad at. Longer generation does not improve character recognition on a faded fax.
The speed gain only shows up on long outputs
At 256 output tokens the two models are identical, 7,229.52 versus 7,229.32 tokens per second. The gap opens as generation continues: at 6,144 output tokens the baseline has decayed to 5,822.87 while Unlimited OCR holds 7,847.71, roughly 35% faster.
On the full OmniDocBench run in base mode at 512 concurrency the advantage shrinks to 12.7% (5,580 against 4,951 tokens per second), because batching already hides much of the per-step attention cost. Single-page invoice processing at high concurrency gains almost nothing from R-SWA.
Accuracy holds to 40 pages, then bends
The paper reports edit distance against page count in a single pass, and the curve is not flat.
Two pages give 0.0362. Ten pages give 0.0526. At 40 or more pages it reaches 0.1069, with Distinct-35 dropping from around 99.9% to 96.90%, which means repeated n-grams start appearing in the output. The 15-page measurement (0.0787) is worse than the 20-page one (0.0572), so treat the curve as a trend rather than a per-page guarantee. Baidu attributes the repetition failures mainly to small text at the 1024×1024 base resolution rather than to attention drift, which matches the tradeoff: multi-page and PDF inputs cannot use the higher-detail crop mode that single images get.
The VRAM math behind "8 GB is enough"
The vLLM recipe states that a single GPU with 8 GB or more is enough for BF16 inference. Community reports do not match that, and the model's own config explains the gap.
The single safetensors file is 6.673 GB. The cache side comes from four fields in config.json: num_hidden_layers: 12, num_attention_heads: 10, num_key_value_heads: 10, v_head_dim: 128, with use_mla: false. Equal query and key/value head counts mean plain MHA, no GQA or MQA sharing to divide by, so cache per token is 2 (K and V) × 12 layers × 10 heads × 128 dims × 2 bytes = 61,440 bytes, or 60 KiB. That gives three figures:
- A full 32K prefill: 32,768 × 60 KiB = 1.875 GiB of cache
- R-SWA decode side, capped by
sliding_window_size: 128: 7.5 MiB, constant - The same decoder without R-SWA at 6,144 output tokens: 360 MiB, growing linearly
These are theoretical cache sizes, not peak allocation: vision encoder activations, allocator fragmentation and the engine's preallocated cache block sit on top. That is why weights plus a long prefill crowd an 8 GB card. One local SGLang run on a 16 GB RTX 4070 Ti Super reported around 12 GB in use, which is consistent with this arithmetic rather than proof of it. Read the 8 GB figure as a short single-page floor, not a spec for the 40-page case.
Those figures also set expectations for what R-SWA buys: capping the decode cache saves hundreds of megabytes rather than gigabytes at this model size, and the gain that shows up in practice is per-step attention cost that stops growing, which is what the throughput curve measures.
There is no official API, so this is how it actually runs
The Hugging Face model page shows "This model isn't deployed by any Inference Provider." There is no first-party endpoint and no Baidu-hosted pricing page for it. Serving it yourself means one of three paths: Transformers with trust_remote_code, SGLang, or the vLLM recipe, which requires vLLM 0.25.0 or newer from the dedicated vllm/vllm-openai:unlimited-ocr container because the architecture is not in a stable pip wheel yet.
Four settings decide whether you get output at all:
- Register the n-gram logits processor (
NGramPerReqLogitsProcessor). Without it, long documents loop on<|det|>coordinate tokens. - Set
ngram_size: 35withwindow_size: 128for single images and1024for multi-page or PDF input. - Start the text content with a literal
<image>token, as in<image>Multi page parsing.No chat template ships with the model. - Pass
skip_special_tokens: False. Leaving the default returns empty strings.
Raw generations carry grounding markup. Keep the text inside <|ref|> and drop the <|det|> bounding boxes to get clean markdown. Page boundaries are not emitted natively either, so ask for page labels in the prompt if you need them for an audit trail.
The recipe's known-good server and request look like this:
docker run --rm --gpus all --network host --ipc host \
vllm/vllm-openai:unlimited-ocr baidu/Unlimited-OCR \
--trust-remote-code \
--logits_processors vllm.model_executor.models.unlimited_ocr:NGramPerReqLogitsProcessor \
--no-enable-prefix-caching --mm-processor-cache-gb 0
client.chat.completions.create(
model="baidu/Unlimited-OCR",
messages=[{"role": "user", "content": [
{"type": "text", "text": "<image>Multi page parsing."},
{"type": "image_url", "image_url": {"url": page_data_url}},
]}],
max_tokens=8192, temperature=0.0,
extra_body={"skip_special_tokens": False,
"vllm_xargs": {"ngram_size": 35, "window_size": 1024}},
)
Use the unlimited-ocr-cu129 image tag on Hopper cards, and note that multiple images in one request fall back to non-crop base mode, which is the case that wants window_size: 1024.
Cost per 1,000 pages: self-host versus a managed API
Open weights are free; running them is not. One of the few published hands-on throughput figures comes from a practitioner in the Hacker News thread who converted roughly 200 pages per hour of a Japanese grammar PDF on an RTX 4090 through Transformers. Put that against rented hardware at RunPod's Community rate of $0.34 per hour for a 4090.
| Option | Cost per 1,000 pages |
|---|---|
| Self-host, single stream (4090 at $0.34/hr, 200 pages/hr) | ~$1.70 |
| Google Enterprise Document OCR, 1K to 5M pages/month | $1.50 (list) |
| Google Layout Parser, same page count | $10.00 (list) |
| Self-host, saturated batch floor (A100 80GB at $1.39/hr, modelled) | ~$0.07 |
List prices as of 2026-07-29. Google's first 1,000 pages per month are free and the rate drops to $0.60 above 5 million pages. The bottom row is a modelled floor rather than a measurement, and it moves with output length: at the paper's 5,580 tokens per second under 512-way concurrency, 700 output tokens per page gives about 28,700 pages per hour ($0.05 per 1,000), 1,000 tokens gives about 20,000 ($0.07), and a dense 2,000-token page gives about 10,000 ($0.14). That throughput was measured on Baidu's own evaluation cluster, not on a rented A100, so the row mixes a benchmark rate with a rental price. Real deployments land above all three once you add idle capacity, retries on failed pages, preprocessing and storage.
A single-stream deployment on a consumer card costs about the same as Google's managed OCR, so self-hosting wins on concurrency and data residency, not on the license. And parsing is only half the job: turning that markdown into fields still means a call to a long-context text model, where you are back to per-token billing rather than per-page, whether that runs on your own stack or through something like the GPT-5.6 API.
Where it breaks
The failure modes people report are the ones a frozen encoder predicts. The same 4070 Ti Super run got garbled output, missed regions and structure drift on receipts, handwriting and complex scans, while clean printed pages came through. Practitioners in the Hacker News thread describe the class of VLM-OCR error that matters for compliance work: foreign words silently translated into English, and a handwritten name "corrected" into its more probable spelling. Both are single anecdotes, so treat them as things to test for rather than as measured rates.
Positioning matters too. Unlimited OCR is not the top of the accuracy tables. On aggregated OmniDocBench listings, PaddleOCR-VL-1.6 sits at 96.33 (vendor self-reported) against 93.92 on v1.6, and the model has not appeared on olmOCR-Bench at all. Per-page accuracy is not the axis it competes on.
Should you run it?
Good fit if your pipeline currently loops over pages and stitches text back together, if your documents are born-digital or cleanly scanned, and if cross-page structures like tables split across a page break are breaking your current output.
Poor fit if your volume is thousands of single-page invoices per day (per-page pipelines batch better and cost less), if handwriting or photographed receipts make up a meaningful share of input, or if you need an SLA and an audit trail today rather than a GPU and a container tag.
Either way, test before committing. A workable minimum: 50 documents from your own corpus, bucketed by length (1 to 5 pages, 6 to 20, 20 or more) and by input quality (born-digital, clean scan, photographed); ground truth typed by hand for 10 of them; then score character error rate, reading-order edit distance and table TEDS separately instead of averaging them, and record wall-clock seconds per page on the GPU you intend to rent. Compare against whatever you run today on the same 50 files, and set the bar where your downstream step actually breaks, which for field extraction usually means table structure rather than raw CER. For reference on how far tuning can move the number, one team running enterprise-scale PDF volume reported 0.94% character error rate after rewriting the inference layer in Rust.
FAQ
Is Unlimited OCR free?
The weights are MIT licensed and free to download from Hugging Face or GitHub, including commercial use. Inference is not free: budget roughly $0.07 to $1.70 per 1,000 pages depending on how well you batch, plus engineering time.
Is there an official Unlimited OCR API?
No. The model page shows no Inference Provider deployment, so any endpoint you find is a third party serving the open weights, priced and rate-limited by that vendor rather than by Baidu.
Is Unlimited OCR the best OCR model right now?
Not on the accuracy tables: PaddleOCR-VL-1.6 reports 96.33 on OmniDocBench against Baidu's 93.92 on v1.6, and the model has no olmOCR-Bench entry yet, so cross-benchmark consistency is unproven. Its measured lead is the 40-page single pass at 0.1069 edit distance.
Can I run Unlimited OCR in Ollama?
The official card documents Transformers, vLLM and SGLang only, and the custom architecture needs trust_remote_code. Community quantizations exist on Hugging Face, but treat any Ollama build as unverified until you have compared its output against the reference path on your own files.