想用文档提取API?先看ExtractBench的溯源评测。多数工具给不出出处,LlamaExtract Agentic Plus在长文档上还挺得住。
ExtractBench对文档提取API的溯源能力进行了严格评测:字段的值及其引用必须同时正确才计分。结果显示,VLM和编码智能体完全不返回证据,两层得分均为零。在能返回框的系统中,最佳词级F1仍低于50%。某专用API在短文档上页面级F1为61.7%,长文档上直接降到0.0%。LlamaExtract Agentic Plus在页面级和词级分别达到84.9%和46.4%,长文档上仍保持87.1%。
Most document extraction APIs can't tell you where a value came from. For ExtractBench, we scored g...
Most document extraction APIs can't tell you where a value came from. For ExtractBench, we scored grounding strictly: a field only counts if the value AND its citation are correct, word-level box at IoU 0.5. A perfect box around a wrong value earns nothing. Results: VLMs and coding agents return no evidence at all — zero at both levels. Among systems that do return boxes, the best word-level F1 is still under 50%. And grounding collapses with length: one specialized API goes from 61.7% page-level on short docs to 0.0% on long ones. LlamaExtract Agentic Plus leads at both levels — 84.9% page-level, 46.4% word-level — and holds at 87.1% on long documents where others hit zero. Every extracted value should come with receipts. ExtractBench now gives the field a baseline to track it 👇 Blog: lnkd.in/gNm97fXp 6 Paper: lnkd.in/euAfScWx F Your browser does not support the video tag. 🔗 View on Twitter 💬 1 🔄 1 ❤️ 4 👀 881 📊 2 ⚡
- Jerry Liu08-16 21:19原文