16 个前沿 VLM 文档解析评测:Opus 5.5 性价比最佳
LlamaIndex 测了 16 个模型读 PDF 的能力,Opus 5.5 表格解析最强还便宜,做智能体文档处理的可以看看这份榜单。
LlamaIndex 对 16 个前沿视觉模型在文档解析任务上做了评测,考察更高推理档位是否带来更好效果。结果显示 Opus 5.5 在价格与效果比上领先,表格解析尤其突出,完整结果发布在 ParseBench 上。Astra 效果也不错但起价更高,GPT-6 Luna 在低价位更有竞争力。大规模文档解析仍推荐 LlamaParse 这类专用 OCR,成本更低、效果更好。
We comprehensively evaluated 16 recent frontier VLMs - including Opus 5.5 and GPT-6 Sol/Luna - on whether higher effort led to better document parsing performance. Higher effort typically leads to improvements on other benchmarks (coding, knowledge work), but up until recently it wasn't obvious that this natively improved capabilities for reading PDFs. Results: ✅ Out of the frontier models, Opus 5.5 has the best performance relative to its price. It's especially good at parsing tables. ✅ Astra is also quite good, but starts at a more expensive price than Opus. ✅ GPT-6 Luna is more compelling at the cheaper end of doc parsing If you're parsing documents at scale, you'll still want a dedicated OCR solution like LlamaParse ( cloud.llamaindex.ai ) that has better performance at a cheaper price. But if you're parsing docs "in the agent loop" within an app like Codex/Claude code, and you're too lazy to integrate a dedicated solution, then Opus 5.5 is currently the leader. Full results on ParseBench: parsebench.ai 💬 3 🔄 1 ❤️ 8 👀 1156 📊 6 ⚡