a16z发了份计算机使用智能体的数据,一年前最强模型才42分,现在到85分了,比人类还高,想了解进展可以看这个。
去年最佳计算机使用模型在真实桌面操作的标准基准上得分42%,如今最佳模型得分已达85%,而人类测试者在相同任务上得分约72%。数据来自a16z分析师Fabrizio Serafini、Seema Amble和Zephratic的调研,显示该领域在过去18个月取得显著进展。
"In the last 18 months, computer-use agents crossed from demo to deployable." A year ago, the best ...
"In the last 18 months, computer-use agents crossed from demo to deployable." A year ago, the best computer-use model scored 42% on the standard benchmark for agents operating a real desktop. Today's best: 85%. Human testers score ~72% on the same tasks. The data on computer-use agents, from @fabrisera2000 , @seema_amble , and @zephratic : a16z.news/p/can-agents-u… Fabrizio Serafini @fabrisera2000 x.com/i/article/2086… 🔗 View Quoted Tweet 💬 0 🔄 3 ❤️ 14 👀 2705 📊 1 ⚡