百度开源Unlimited OCR模型
百度开源的轻量级OCR模型,能一次性处理100页PDF,完全本地运行,免费且开源。
Unlimited OCR是百度推出的开源OCR模型,仅3B参数。在ParseBench基准测试中得分46.2%,排名第115。相比PaddleOCR,表格识别略优,布局识别略差。支持32K上下文窗口,可一次性处理100页PDF文档,错误率低于0.11。
Unlimited OCR is a good example of a low-parameter OCR model that can process any document extremely quickly. It's good at reading order and layout, and decent over tables. It is slightly better than PaddleOCR on tables and slightly worse on layout. There are tradeoffs though - it ignores visual elements, formatting, and the most complex tables. It scores ~46.2% and ranks #115 on ParseBench. There's plenty of opportunities to improve this frontier, but Unlimited OCR is still a good milestone! Check out ParseBench: parsebench.ai Brooks Whale X 🐋 @BrooksWhaleX 🚨 CHINA JUST OPEN SOURCED AN OCR THAT EATS 100 PAGE PDFS IN ONE PASS It’s called Unlimited OCR. Only 3B params. Runs locally. Every other OCR tool chops your doc into pages and loses the thread. This one reads the whole thing in one continuous pass. → One shot full context parsing (32K context window) → Multilingual out of the box → 93% on the standard parsing benchmark (6 points over baseline) → Under 0.11 error rate past 40 pages → Runs 100% locally on your own hardware → Works with Transformers, vLLM, SGLang, Docker, Ollama, llama.cpp Traditional cloud OCR (Textract, Google Vision, Azure Doc Intelligence) costs $1.50 to $15 per 1,000 pages. This runs on your machine. For free. Forever. Baidu built it explicitly to push DeepSeek OCR one step further. Already at 1.9M downloads on Hugging Face and most people have no idea it exists yet. 100% open source. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 7 🔄 2 ❤️ 18 👀 1555 📊 10 ⚡