Liteparse 升级为最快 PDF 解析器,支持边界框输出

Last week we revamped Liteparse to be the fastest PDF parser out there ⚡️ An underrated part of lit...

精选理由

做文档解析或 AI 代理的开发者终于有了一个又快又准的开源选择——Liteparse 的边界框输出让审计追踪变得简单,值得直接试。

AI 摘要

LlamaIndex 创始人 Jerry Liu 宣布 Liteparse 完成重大升级,成为目前最快的 PDF 解析器。新版用 Rust 重写了整个库,并适配为 Python 和 Node 原生包,支持 50 多种文档类型。除了提取文本,Liteparse 还能输出边界框,让编码代理可以精确追溯源文档。团队正在开发 Markdown 支持,并鼓励用户提交 issue 和 PR。

原文 · Jerry Liu

Last week we revamped Liteparse to be the fastest PDF parser out there ⚡️ An underrated part of lit...

Last week we revamped Liteparse to be the fastest PDF parser out there ⚡️ An underrated part of liteparse is it doesn't just give you text. It gives you bounding boxes that a coding agent can use to paint exact audit trails back to the source document. For instance, check out the deep research skill we compiled in liteparse_samples: github.com/jerryjliu/lite… Come check out liteparse: github.com/run-llama/lite… We are hard at work making liteparse even better (e.g. Markdown support). Please feel free to open up issues, PRs, and let us know your feature requests 🙏 Your browser does not support the video tag. 🔗 View on Twitter Jerry Liu @jerryjliu0 We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf, pypdf, markitdown, pdftotext, opendataloader, pymupdf4llm) Introducing LiteParse v2 - we rewrote the entire library into Rust and adapted it as native packages for Python and Node. It supports 50+ different document types, can be triggered directly or installable directly within your favorite AI agent. Blog: llamaindex.ai/blog/liteparse… Repo: github.com/run-llama/lite… 🔗 View Quoted Tweet 💬 2 🔄 5 ❤️ 10 👀 2274 📊 7 ⚡

  • AI Notkilleveryone05-31 16:24原文
  • IT之家06-02 16:58原文