LlamaIndex 提出 Agentic OCR:用多轮校验替代单次识别
LlamaIndex 的 Logan 写了篇实操文:OCR 别再单次跑,改成布局感知、难元素路由、多轮自检的循环,能救回表格和图表。
LlamaIndex 开源负责人 Logan Markewich 撰文介绍 Agentic OCR 的工作方式。传统 OCR 只做单次识别,表格会被压平、图表丢失、多栏排版顺序混乱,且没有机制检查输出。Agentic OCR 把解析变成一个循环:按版面感知的阅读顺序解析,把复杂元素路由给合适的模型,再做多轮校验和自我纠错。文章指出前沿模型在成本和延迟上对 OCR 这类任务是过度配置,但在大量复杂边缘场景中仍会出错。
OCR is the perfect example of a use case that has been dominated by legacy/brittle systems, and can be solved both accurately and cheaply by applying the right amount of agentic intelligence. A properly tuned agentic OCR needs to dynamically apply extra compute to complex elements, review and correct failures, and create the right semantic meaning throughout the page. It's a task that frontier models are both overengineered for in terms of cost and latency, and struggle at over a long tail of complex edge cases. Check out this blog post by @LoganMarkewich below: llamaindex.ai/blog/ocr-is-de… LlamaIndex 🦙 @llama_index OCR is dead 🪦 Long live agentic OCR! Traditional OCR makes one pass and hands back whatever text it got. Tables get flattened, charts disappear, and multi-column layouts come out scrambled. Nothing checks the output. Agentic OCR treats parsing as a loop instead: ✅ layout-aware reading order ✅ smart routing of hard elements to the right model ✅ multi-pass verification and self-correction ✅ multimodal parsing of charts, images, and complex tables Logan Markewich, LlamaIndex's Head of Open Source, wrote about what that shift looks like in practice and where it's still struggling. Link to the breakdown in the comments below. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 2 👀 558 📊 1 ⚡