ExtractBench上线测试前沿模型文档提取能力

We've benchmarked all the latest frontier models on hard document extraction tasks (with Fable 5.1 a...

精选理由

Llama团队发布了ExtractBench基准,测试了5.6 Sol等模型在复杂文档提取上的表现,还对比了LlamaParse等工具。

AI 摘要

ExtractBench已在Kaggle上线,测试了5.6 Sol、OpenAI模型、Gemini 3 Flash和Opus 5等前沿模型在文档提取任务中的表现。该基准测试涵盖370个企业文档,涉及8个商业领域和67种文档类型,包括长记录列表、噪声扫描、手写和复杂表格等。LlamaParse作为专用提取工具,在价格/性能前沿提供基础引用功能。

原文 · Jerry Liu

We've benchmarked all the latest frontier models on hard document extraction tasks (with Fable 5.1 a...

We've benchmarked all the latest frontier models on hard document extraction tasks (with Fable 5.1 and Gemini 3.8 Flash coming soon!). ExtractBench is now live on @kaggle . 5.6 Sol leads the pack, followed by other OpenAI models, then Gemini 3 Flash and Opus 5. Come check it out! kaggle.com/benchmarks/lla… If you're looking for a dedicated extraction tool with out of the box grounding, citations at the price/performance frontier, come check out LlamaParse: cloud.llamaindex.ai LlamaIndex 🦙 @llama_index ExtractBench is now live on @Kaggle . It tests schema-guided document extraction on the documents most likely to break downstream agents and workflows, including long record lists, noisy scans, handwriting, and complex tables. The benchmark covers 370 enterprise documents across 8 business domains and 67 document types. See how LlamaParse, Codex, Claude Code, and more stack up. Read the full story → llamaindex.ai/blog/llamainde… 🔗 View Quoted Tweet 💬 2 🔄 2 ❤️ 8 👀 1825 📊 3 ⚡