Skip to main content

Baidu's Unlimited-OCR: The AI That Reads an Entire Book in One Pass

Baidu Unlimited-OCR

On June 22, 2026, Baidu quietly open-sourced a model that solves one of OCR's most stubborn problems: reading long documents. Called Unlimited-OCR, it can parse a 40-page contract, a scanned textbook, or a dense financial report in a single forward pass — and it does so while keeping memory dead flat. The GitHub repo crossed 10,000 stars within five days. Here's why it matters, in plain terms.

The problem: OCR forgets how to breathe on long documents

Modern OCR is done by vision-language models that "read" a page image and emit text. The catch is the KV cache — the running memory a transformer keeps of everything it has seen so far. On a one-page receipt that's trivial. On a 40-page document, that cache balloons: memory climbs, per-page latency creeps up, and eventually the model runs out of room or slows to a crawl. In practice, most OCR systems quietly punt by slicing a document into pages and stitching the results back together — which loses cross-page context (think a table that spans three pages, or a footnote referenced ten pages later).

Unlimited-OCR's whole pitch is that it doesn't slice. It reads the whole thing at once, and its memory footprint stops growing.

Unlimited-OCR keeps the KV cache flat as documents get longer

How they pulled it off

Two ideas do the heavy lifting. The first is the DeepEncoder, inherited from the DeepSeek-OCR lineage Unlimited-OCR is built on. It compresses a page hard — a 1024×1024 image becomes just 256 visual tokens, a roughly 16× squeeze. Fewer tokens per page means more pages fit inside the model's 32K-token context window.

The second, and the real innovation, is Reference Sliding Window Attention (R-SWA). Instead of letting the decoder attend to (and cache) every token it has ever produced, R-SWA bounds the cache to a fixed size — roughly the current page's tokens plus a small window of recent output (128 tokens by default). The result: as the document gets longer, memory stays flat and per-step latency stays flat. That's the green line in the chart above — and it's the difference between "reads 2 pages" and "reads 40+."

Under the hood it's a Mixture-of-Experts (MoE) model: 3 billion parameters total, but only about 500 million active per token. You get the capacity of a bigger model at the cost of a much smaller one — which is exactly what you want for something meant to grind through long documents cheaply.

Does it actually work?

The numbers say yes. On OmniDocBench, the standard end-to-end document-parsing benchmark, Unlimited-OCR posts a comprehensive 93.23 on v1.5 — beating DeepSeek-OCR by 6.22 points — and 93.92 on v1.6, a new state-of-the-art in the paper. It's also faster: about 5,580 tokens/second versus DeepSeek-OCR's 4,951, a 12.7% bump. And on 40-plus-page documents, its edit distance stays below 0.11 — meaning the transcription barely degrades as the page count climbs, which is precisely where older systems fall apart.

Two resolution modes cover the range: a Base mode (1024×1024) for dense single pages, and a dynamic Gundam mode for multi-page batches. It's multilingual, ships in the usual Transformers/Safetensors format, and — the part that will get it adopted — it's released under a permissive MIT license with open weights on GitHub and Hugging Face (baidu/Unlimited-OCR).

Why it's a big deal

OCR has been "solved" for short text for years; the frontier has moved to documents as documents — contracts, research papers, medical records, filings — where structure and cross-page context are the whole point. Feeding those into a RAG pipeline or an LLM has meant brittle page-splitting and a lot of glue code. A model that ingests the entire document in one shot, keeps the layout coherent, and runs on a modest activated-parameter budget removes a real bottleneck for anyone building document AI.

The honest caveat: benchmark leadership is a snapshot, not a moat — DeepSeek and others will answer, and OmniDocBench scores don't capture every messy real-world scan. But the architectural idea here — bounding the KV cache so cost stops scaling with length — is the kind of thing that outlives any single leaderboard. For a free, MIT-licensed model you can run today, "reads the whole book at once" is a genuinely new capability, not a press release.

If you work with long PDFs, it's worth a look: github.com/baidu/Unlimited-OCR.


🔗 Explore more from Syncster

Comments

Popular posts from this blog

Cursor AI Review: Is the AI Code Editor Worth It?

I've been using Cursor as my main code editor for a while now, and enough people have asked whether it's worth switching to that a proper review felt overdue. Short version: for me, yes — but with caveats. What is Cursor? Cursor is an AI-first code editor built as a fork of VS Code. That means every extension, theme, and keybinding you already use in VS Code works here, but with AI woven directly into the editing experience instead of bolted on as a plugin. It's made by Anysphere and can run models from OpenAI and Anthropic under the hood. What I like Tab completion is uncanny. Cursor predicts your next edit — not just the rest of the line, but the next change across the file. Once you get used to hitting Tab, going back to a plain editor feels slow. The Composer / Agent mode. You describe a change in plain language and it edits multiple files at once, showing you a diff to accept or reject. For refactors and boilerplate, this saves real time. It unde...

MacBook Pro M5 vs M5 Pro: Which One Should You Actually Buy?

Apple's latest 14-inch MacBook Pro comes in two very different flavors: the base M5 and the step-up M5 Pro . On paper they look similar — same gorgeous Liquid Retina XDR display, same design — but under the hood the gap is bigger than the names suggest. Here's a clear, no-hype breakdown, with concrete use cases so you can match the chip to your work. Quick spec comparison Spec M5 M5 Pro CPU 10-core (4 performance + 6 efficiency) Up to 18-core (6 performance + 12 efficiency) GPU 10-core Up to 20-core Neural Engine 16-core 16-core Memory bandwidth 153 GB/s 307 GB/s (roughly double) Unified memory 16 / 24 / 32 GB 24 / 48 / 64 GB Max storage Up to 4 TB SSD Up to 8 TB SSD Battery (video playback) Up to 24 hours Up to 22 hours Media engines Single encode/ProRes engine More encode/ProRes engines (higher configs) What actually changes between them More cores — the M5 Pro nearly doubles CPU cores and adds GPU cores, so sustained, multi-threaded work finishe...

How I used Google Sheets and Apps Script

Google Sheet is one of the most powerful spreadsheet application that exists online, rivaling with Microsoft's Excel. One of the main strengths is its strong support for collaboration with other users, much easier and popular than collaboration tools with Microsoft Office. Aside from plain spreadsheet, it also supports extensions such as macro. If you are familiar with macros on other office tools, they work almost the same. However, the most extension I use and tinker with is the Apps Scipt . Apps Script Extension One of the challenges I faced recently is how do I track or monitor reports in our department if they are submitted on time or worst, forgotten due to lack of better monitoring tools. So I thought if there can be simple applications that can be deployed or use by a more general user to allow reminding periodically what reports are approaching due dates or those that are past dues. Then I looked for a way, instead of creating a full blown app from scratc...