LiteParse能直接读PDF勾选框和表单字段,毫秒级,还告诉你何时该用LlamaParse。
LlamaIndex发布LiteParse更新,无需视觉模型即可从PDF中提取表单字段值、勾选状态、注释、嵌入图像、矢量图形等结构化数据。处理速度达到每页毫秒级。对于需要模型处理的页面,LiteParse会输出复杂度信号,标记扫描页、多栏文本、表格和密集图形,帮助用户把解析任务路由到LlamaParse等工具。
You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structure...
You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form field values, checkbox states, annotations, embedded images, vector graphics, tagged document structure, and word-level bounding boxes, all running in ms per page. For pages that do need a model, new complexity signals tell you why. Scanned pages, multi-column text, tables (ruled and borderless), and dense figures help you route your parsing to the best tools (like LlamaParse!) Docs 👉️ : developers.llamaindex.ai/liteparse/guid… U developers.llamaindex.ai/liteparse/guid… V 👩💻 Try i github.com/run-llama/pars… 8qn 💬 0 🔄 0 ❤️ 2 👀 516 ⚡