ExtractBench最新测试了20多个模型,Qwen 3.8模型表现抢眼,新版本GLM-5.3-flash和qwen 3.8 flash也值得关注。
ExtractBench测试20多个模型在文档提取任务上的表现,Qwen 3.8模型表现最佳,GLM-5.3-flash和qwen 3.8 flash新版本也刚刚发布。测试结果基于统一F1的“价值准确性”指标,不包含视觉定位和置信度评分。
We benchmarked 20+ open-weight models on easy-to-hard document extraction tasks through ExtractBench...
We benchmarked 20+ open-weight models on easy-to-hard document extraction tasks through ExtractBench. The results are all available on @huggingface 🤗 ExtractBench is a schema-guided extraction benchmark that contains 4.8k+ pages across 8 domains and 67 document types, with a mix of short, medium, long docs and simple/complex schema.s The results reported is a measure of “value accuracy” through unified F1. ✅ Qwen 3.8 leads the pack ✅ earlier generations of Qwen models are also quite strong ✅ kimi-k3 is the next best. GLM-5.3-flash and qwen 3.8 flash also just released today - hopefully will have results on these soon! Note : these results don’t include visual grounding (whether each value is mapped to the right bounding box) and confidence scores. When you include these, it is increasingly clear why specialized OCR tools (like LlamaParse) matter, since this is metadata that’s hard to DIY by prompting the raw model. Come check out our HF leaderboard: huggingface.co/datasets/llama… U Learn more about ExtractBench here: extractbench.ai B 💬 2 🔄 0 ❤️ 4 👀 378 📊 3 ⚡