批量解析复杂企业文档?LlamaExtract比Claude Code便宜一半还更准,去LlamaParse体验。
LlamaIndex推出文档抽取引擎LlamaExtract Agentic Plus,在ExtractBench上达到95.6%的value accuracy。ExtractBench包含370份企业文档、4869页、67种文档类型,覆盖财务、能源、政府等领域。该基准发现,超过50页的长文档会让商业视觉语言模型的召回率跌破35%,原因是静默截断列表。LlamaExtract Agentic Plus定价约为Claude Code/Codex的25%-50%,每页成本低于最接近竞品的三分之一。
This week we launched the world's most accurate document extraction agent over real-world documents....
This week we launched the world's most accurate document extraction agent over real-world documents. Introducing LlamaExtract Agentic Plus 💫 . It is a complete document extraction model+harness engine that can convert even the most complex document types/lengths/schemas into clean, grounded, structured output. We benchmarked it over ExtractBench - a comprehensive set of real-world enterprise documents. It is measured over various dimensions: document length (short/medium/long), task (long-list, needle in haystack, dense docs), perception (rotated, scanned, handwriting), table structured (merged headers, cross-page tables, table within cells), and business domains. We compared it against a variety of coding agent harness to document extraction solutions. LlamaExtract Agentic Plus is not only highly accurate, it is quite price efficient relative to its capabilities (25-50% the price of Claude Code/Codex). You can see the full release of ExtractBench here: llamaindex.ai/blog/introduci… z If you want to check it out, come sign up for an account on LlamaParse: cloud.llamaindex.ai A Jerry Liu @jerryjliu0 Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A well-tuned extractor must parse multi-page filings without dropping rows, emit exact spatial citations for auditability, and handle messy scans. Also they must do all of this at a viable per-page cost so that you can scale this to millions of docs in production (you can’t be paying upwards of $1 in tokens per page!) Existing extraction benchmarks fall short: they are not large/diverse enough in document domain (finance, energy, gov, auto), elements (long records, scans, grounding), and schemas. So our applied research team built ExtractBench. We evaluated 14 systems: frontier VLMs, coding agents, and specialized extraction APIs, against 370 enterprise documents: 4,869 pages, 67 document types. Our biggest finding 🧪: Short documents mask critical system flaws. On files past 50 pages, commercial VLMs collapse below 35% recall due to silent list truncation. They hold high precision, but lose output attention and drop most of the table rows. ExtractBench evaluates value accuracy, long-record completeness, spatial grounding, and per-page cost with zero LLM judges. It is 100% deterministic and reproducible. In tandem with ExtractBench, we’re also introducing 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗣𝗹𝘂𝘀, a new Extract tier in LlamaParse that debuts at #1 on the leaderboard: 95.6% value accuracy, at less than a third the cost of the closest peer. Explore the findings, download the dataset, or run the harne llamaindex.ai/blog/introduci… o/tsXfovK github.com/run-llama/Extr… o/auzmOGgUl3 H huggingface.co/datasets/llama… o/1eKisSBrE8 We will be actively evolving both our extraction benchmark as well as our extraction harness over time. If you check out either ExtractBench or LlamaParse, let us know your feedback! Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 9 👀 1153 📊 3 ⚡
- LlamaIndex14:05原文