这篇论文戳破了AI基准测试的泡沫——高分不等于能干实事。做AI自动化部署的团队、评估智能体能力的开发者,看完会重新审视自己的测试标准,建议点开看看真实工作场景的差距。
一篇新论文提出“Agents' Last Exam”基准测试,要求AI智能体完成来自55个数字工作领域的真实专家任务,包括工程、金融、医学、法律、媒体和科学。测试发现,当前最强的智能体系统在最难任务上的平均完全通过率仅为2.6%,远低于其基准分数所暗示的水平。该基准强调从“能否回答难题”转向“能否完成人们付费做的工作”,使用自动检查或严格评分标准而非主观评判。结果表明,基准测试的成功尚未转化为广泛的工作场所能力,智能体在真实自动化中仍不可靠。
Today’s frontier agents are far less ready for rea…
Today’s frontier agents are far less ready for real-world automation than their benchmark scores suggest.
This paper proposes a Agents’ Last Exam, a benchmark that asks AI agents to finish real expert work, and today’s agents mostly fail.
Even strong agents of today are nowhere near reliable on the hardest real workflows, which means benchmark success has not yet become broad workplace capability.
So this paper shifts the question from “can AI answer hard questions?” to “can AI complete real work that people get paid to do?”
Most of today's AI benchmarks show impressive scores, but they do not prove that agents can finish useful work in real jobs.
Agents’ Last Exam tries to fix this by testing agents on long tasks from 55 digital work areas, including engineering, finance, medicine, law, media, and science.
The tasks come from experts’ real completed projects, and the agent must use normal computer tools like files, browsers, command lines, and desktop software to produce a finished result.
The authors tested many current agent systems and models, then scored their finished work with automatic checks or strict rubrics instead of loose human opinions.
The main result is that today’s best systems still struggle badly, with an average full pass rate of only 2.6% on the hardest tier.
----
Link – arxiv. org/abs/2606.05405
Title: "Agents' Last Exam"