AI产品精选

LlamaIndex:前沿模型并未让文档OCR商品化

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whe...

精选理由

LlamaIndex用数据反驳“OCR只是功能”,说LlamaParse比GPT等前沿模型更准更便宜,做文档解析的可以看看。

AI 摘要

LlamaIndex发布博客称,前沿模型在文档理解性能上趋于平缓,GPT-5.5到5.6、Gemini 3.5到3.6、Opus 4.8到5的视觉理解基准无明显提升。跨三个GPT世代,解析准确率仅提升约24点,而每页成本翻了4倍。LlamaParse通过代理模式和低成本模式,在密集表格和图表等复杂边界情况上超越了前沿模型。该公司表示,可将前沿模型蒸馏为参数效率更高的模型,以更低成本实现同等性能。

原文 · Jerry Liu

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whe...

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all document processing solutions - just screenshot the page and feed it to your favorite frontier model. 1️⃣ Frontier models are flatlining in document understanding performance. Incremental releases in every model version (gpt 5.5 -> 5.6 sol, Gemini 3.5 flash -> 3.6 flash, opus 4.8 -> opus 5) are not improving visual understanding benchmarks. 2️⃣ The Pareto frontier for document OCR is much higher than the frontier models, and will always remain much higher. We’ve carefully tuned our agentic and cost-effective modes to solve a long tail of complex edge cases (dense tables, charts) that frontier models don’t care about. Also for equivalent performance, there’s always ways to get much lower cost. 3️⃣ Even if they were getting better, you can distill / posttrain them into much more parameter efficient models for a fraction of the cost. Different document pages can be routed to different processors of varying complexity. 4️⃣ Every startup is doing the same thing right now. Focusing on one task means you can always hillclimb that task more effectively than a general model over the task, in terms of accuracy/cost/latency. Check out the blog: llamaindex.ai/blog/document-… LlamaParse has gotten a LOT better in the past few months. Come check it out! cloud.llamaindex.ai LlamaIndex 🦙 @llama_index "OCR is just a feature now. Frontier models will eat it." We hear this constantly. The data says otherwise. Across three GPT generations, parsing accuracy gained ~24 points, while cost per page 4x'd. And the newest frontier models still trail specialized parsers. Read more below on why the gap persists (and compounds) 👇️ llamaindex.ai/blog/document-… b 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 262 ⚡