Opus 5 与 GPT-5.6 Sol 在浏览器任务上表现接近

Opus 5 and GPT-5.6 Sol are neck-and-neck on this!

精选理由

Opus 5 和 GPT-5.6 Sol 在一个新浏览器基准上打平了,这个基准是真实用户任务,评估标准很细,看看谁更会上网吧。

AI 摘要

Opus 5 和 GPT-5.6 Sol 在浏览器使用基准测试中表现接近。该基准由数百名用户参与,将真实用户任务转化为测试。测试采用原子化、可验证且明确的评估标准。这是目前最好的浏览器智能体基准测试。

原文 · Browser Use

Opus 5 and GPT-5.6 Sol are neck-and-neck on this!

Opus 5 and GPT-5.6 Sol are neck-and-neck on this! Alexander Yue @Alezander9 We talked with hundreds of users, and turned the hardest real user tasks into a browser-use benchmark We invested heavily into atomic, verified, unambiguous rubrics for LLM judges. This is the best browser agent benchmark ever created 🧵 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 8 👀 900 📊 1 ⚡