AI模型精选

ExtractBench:14个文档提取系统各有感知盲区

Every document extraction system has a perception blind spot. We mapped them. For ExtractBench, we ...

精选理由

LlamaIndex测了14个文档提取系统,Codex、Gemini在扫描件上都掉链子,只有Agentic Plus三项全在93%以上。

AI 摘要

ExtractBench测试了14个文档提取系统,使用1950年代监管文件、手填税表和手机拍摄等非数字化文档。Codex在扫描和手写场景准确率超93%,但处理旋转或纯图像页面时降至约80%。专用API表现相反,旋转和手写识别良好,扫描准确率为81%。Gemini 3.5 Flash从干净PDF的88.6%跌至扫描件的71.1%。LlamaIndex的Agentic Plus是唯一无盲区系统,在旋转、扫描和手写三项上分别达到95.9%、93.9%和93.8%。

图片来源 · LlamaIndex
原文 · LlamaIndex

Every document extraction system has a perception blind spot. We mapped them. For ExtractBench, we ...

Every document extraction system has a perception blind spot. We mapped them. For ExtractBench, we tested 14 systems on documents that weren't born digital: 1950s regulatory filings, hand-filled tax forms, and pages degraded with fax thresholding, photocopier tone curves, sensor noise, and phone-camera capture. The failures don't overlap. 💡 Codex reads scans and handwriting above 93%, then drops to ~80% on rotated or image-only pages. 💡 Specialized APIs are the exact inverse: fine on rotation and handwriting, 81% on scans 💡 Gemini 3.5 Flash falls from 88.6% to 71.1% the moment a page is scanned. You benchmark on clean PDFs. Production sends you a shadowed photocopy from 1953. Our new Extract Tier, Agentic Plus, was the only system with no blind spot: 95.9% / 93.9% / 93.8% across rotated, scanned, and handwritten, a 2-point spread where others swing 10+. Learn more about ExtractBench 👇️ Bl lnkd.in/gNm97fXp ceB6 Pap lnkd.in/euAfScWx kQpF Your browser does not support the video tag. 🔗 View on Twitter 💬 1 🔄 2 ❤️ 4 👀 687 📊 2 ⚡