AI模型精选

LlamaIndex发布ExtractBench:企业文档信息抽取基准

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction fr...

精选理由

做文档抽取的团队看过来,LlamaIndex 出了个新基准,测了14个系统在370份企业文档上的表现,发现超过50页商业模型召回率就崩了,赶紧看看你的抽取方案中招没。

AI 摘要

LlamaIndex应用研究团队推出ExtractBench,号称最全面的企业复杂文档信息抽取基准。该基准测试了14个系统,包括前沿VLM、编程智能体和抽取API,覆盖370份企业文档、4869页和67种文档类型。评测完全确定性,不使用LLM作为裁判。最大发现是超过50页后,商业VLM的召回率跌破35%,精度虽高但会静默丢失大部分表格行。

图片来源 · LlamaIndex
原文 · LlamaIndex

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction fr...

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems — frontier VLMs, coding agents, extraction APIs — on 370 enterprise docs, 4,869 pages, 67 doc types. Zero LLM judges, fully deterministic. Biggest finding: past 50 pages, commercial VLMs collapse below 35% recall. Precision stays high, but they silently drop most of the table rows. What is your extraction agent missing? Run ExtractBench to see tod llamaindex.ai/blog/introduci… o/5UE8frC github.com/run-llama/Extr… o/27vJ36bL54 H huggingface.co/datasets/llama… o/oeGlwPjcQh Your browser does not support the video tag. 🔗 View on Twitter 💬 2 🔄 1 ❤️ 10 👀 698 📊 3 ⚡

LlamaIndex发布ExtractBench:企业文档信息抽取基准 · AI 热点