OpenAI发布GPT 5.6,ParseBench测试文档理解

@OpenAI released GPT 5.6 today and we ran a day 0 benchmark in ParseBench to test improvements in d...

精选理由

OpenAI刚发了GPT 5.6,Llama Index实测:读文字表格很强,看图表还差点意思。便宜6倍的Luna版性价比很香。

AI 摘要

OpenAI今日发布GPT 5.6系列模型。Llama Index在ParseBench基准上进行测试,评估文档理解能力。新模型在阅读文本和表格方面表现优异,但在图表和布局上仍有困难。Luna版本比Sol版本便宜约6倍,但性能下降微小,表明增加推理token不一定改善视觉理解。

原文 · LlamaIndex

@OpenAI released GPT 5.6 today and we ran a day 0 benchmark in ParseBench to test improvements in d...

@OpenAI released GPT 5.6 today and we ran a day 0 benchmark in ParseBench to test improvements in document understanding. The new family of models continues to excel at reading text and tables, but continues to struggle with charts and layout. What's most interesting is that Luna is about 6 x cheaper than Sol and only results in minor degradations across all ParseBench metrics, showing that an increase in reasoning tokens doesn't always result in a commensurate improvement in visual understanding. 💬 2 🔄 1 ❤️ 13 👀 9185 📊 4 ⚡

  • Jerry Liu07-09 23:07原文
  • SuperTechFans07-11 00:01原文
  • Greg Brockman07-09 17:26原文
  • Simon Willison’s Weblog07-09 19:46原文
  • marktechpost07-09 20:45原文
  • @koltregaskes07-08 06:02原文
  • Sam Altman07-09 02:42原文
  • OpenAI Blog07-09 10:00原文
  • The Rundown AI07-09 17:38原文
  • @OpenAIDevs07-09 17:41原文