技巧

LlamaIndex 详解 Agentic OCR:把文档解析做成多轮自校验循环

精选理由

LlamaIndex 的 Logan Markewich 写了篇文档解析新玩法,传统 OCR 一遍过就完事,现在改成循环识别加自检,表格图表都能抠出来,做 RAG 的可以看看。

LlamaIndex 开源负责人 Logan Markewich 撰文对比传统 OCR 与 agentic OCR 的差异。传统 OCR 只做单次识别,表格被压平、图表丢失、多栏排版错乱,且输出无人校验。Agentic OCR 把解析改成循环流程,包含版面感知的阅读顺序、难解析元素的智能路由、多轮校验与自纠错,以及对图表、图片和复杂表格的多模态解析。文章同时说明了这套流程目前的落地情况与仍在挣扎的环节。

原文 · LlamaIndex

OCR is dead 🪦 Long live agentic OCR! Traditional OCR makes one pass and hands back whatever text it got. Tables get flattened, charts disappear, and multi-column layouts come out scrambled. Nothing checks the output. Agentic OCR treats parsing as a loop instead: ✅ layout-aware reading order ✅ smart routing of hard elements to the right model ✅ multi-pass verification and self-correction ✅ multimodal parsing of charts, images, and complex tables Logan Markewich, LlamaIndex's Head of Open Source, wrote about what that shift looks like in practice and where it's still struggling. Link to the breakdown in the comments below. 💬 4 🔄 1 ❤️ 38 👀 2427 📊 11 ⚡