AI模型精选

GPT-5.6 文档理解基准测试:与上一代无显著变化

We comprehensively benchmarked GPT-5.6 on document understanding. At a high-level there's no change...

精选理由

OpenAI 新 GPT-5.6 文档理解跑分出炉,Sol 没提升但 Luna 便宜六倍,值得看具体差距。

AI 摘要

OpenAI 发布 GPT-5.6 Sol 和 Luna 两个版本。在 ParseBench 上全面测试文档理解,GPT-5.6 Sol 在表格、文本、图表、布局等指标上与 GPT-5.5 性能持平。模型擅长读取文本和表格,但依旧在图表转录和布局边界框提取上表现不佳。Luna 成本仅为 Sol 的 1/6,在所有度量上只有轻微下降。

原文 · Jerry Liu

We comprehensively benchmarked GPT-5.6 on document understanding. At a high-level there's no change...

We comprehensively benchmarked GPT-5.6 on document understanding. At a high-level there's no change between GPT-5.6 Sol and GPT-5.5 in terms of performance over tables, text, charts, layout, and more. The GPT-class of models typically does quite well over table understanding, but they struggle with transcribing complex text layouts and formatting, with transcribing charts, and with deriving bounding boxes over source elements. Come check out our leaderboard of over 70+ frontier models, open-weight models, and OCR solutions on ParseBench: parsebench.ai LlamaIndex 🦙 @llama_index @OpenAI released GPT 5.6 today and we ran a day 0 benchmark in ParseBench to test improvements in document understanding. The new family of models continues to excel at reading text and tables, but continues to struggle with charts and layout. What's most interesting is that Luna is about 6 x cheaper than Sol and only results in minor degradations across all ParseBench metrics, showing that an increase in reasoning tokens doesn't always result in a commensurate improvement in visual understanding. 🔗 View Quoted Tweet 💬 4 🔄 1 ❤️ 6 👀 899 📊 4 ⚡

GPT-5.6 文档理解基准测试:与上一代无显著变化 · AI 热点