LlamaParse新出的Agentic Plus能一次抽几万字段,长文档不丢行,成本还比同类低一大截。
LlamaParse推出新的Agentic Plus抽取模式,面向100-500页长文档,可提取10k-100k个字段,并在FTX债权人矩阵测试中处理了75k字段、114页文档。配套发布的ExtractBench基准覆盖370份企业文档、4,869页、67种类型,评估了14套系统。测试显示,超过50页后商用VLM的召回率低于35%,主要因为静默截断列表。Agentic Plus在该基准上以95.6%的值准确率排名第一,成本不到最接近对手的三分之一。
Our "agentic plus" extractor in LlamaParse is great for extracting out massive volumes of fields (e....
Our "agentic plus" extractor in LlamaParse is great for extracting out massive volumes of fields (e.g. 10k-100k+ fields) from long documents (100-500 pages) We tested with a doc that contains a giant matrix of all creditors for FTX 💸 (75k fields, 114 pages). See screenshot below. You can find full benchmark results on ExtractBench extractbench.ai r). In the meantime, come try out the mode in our playground! cloud.llamaindex.ai A Jerry Liu @jerryjliu0 Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A well-tuned extractor must parse multi-page filings without dropping rows, emit exact spatial citations for auditability, and handle messy scans. Also they must do all of this at a viable per-page cost so that you can scale this to millions of docs in production (you can’t be paying upwards of $1 in tokens per page!) Existing extraction benchmarks fall short: they are not large/diverse enough in document domain (finance, energy, gov, auto), elements (long records, scans, grounding), and schemas. So our applied research team built ExtractBench. We evaluated 14 systems: frontier VLMs, coding agents, and specialized extraction APIs, against 370 enterprise documents: 4,869 pages, 67 document types. Our biggest finding 🧪: Short documents mask critical system flaws. On files past 50 pages, commercial VLMs collapse below 35% recall due to silent list truncation. They hold high precision, but lose output attention and drop most of the table rows. ExtractBench evaluates value accuracy, long-record completeness, spatial grounding, and per-page cost with zero LLM judges. It is 100% deterministic and reproducible. In tandem with ExtractBench, we’re also introducing 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗣𝗹𝘂𝘀, a new Extract tier in LlamaParse that debuts at #1 on the leaderboard: 95.6% value accuracy, at less than a third the cost of the closest peer. Explore the findings, download the dataset, or run the harne llamaindex.ai/blog/introduci… o/tsXfovK github.com/run-llama/Extr… o/auzmOGgUl3 H huggingface.co/datasets/llama… o/1eKisSBrE8 We will be actively evolving both our extraction benchmark as well as our extraction harness over time. If you check out either ExtractBench or LlamaParse, let us know your feedback! Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 330 ⚡