论文多源确认精选73°

Agent基准测试需评估评分者而非仅任务

精选理由

Zapier的AutomationBench基准测试被多家前沿实验室引用,但评分者存在严重缺陷,修复后评分差异达27.9%。

Parsewave审查了Zapier的AutomationBench中全部600个公开任务,发现206个评分者存在缺陷。研究人员使用智能体生成看似正确但实际错误的答案,验证评分者准确性。修复后的评分者在1,235次Kimi K3测试中,27.9%的情况下给出不同评分结果。任务813案例显示,原评分者仅检查Salesforce笔记而忽略DocuSign模板,错误通过13项检查。

原文 · rohanpaul_ai

Agent benchmarks need audits of the graders, not only the tasks.

Parsewave reviewed all 600 public tasks in Zapier's AutomationBench, which frontier labs cite on their model cards.

It used agents to write convincing wrong answers, then had humans confirm which graders actually failed: 206 did.

AutomationBench Verified fixes every one.

On 1,235 Kimi K3 runs, the fixed graders gave a different verdict 27.9% of the time.

The clearest case is task 813. The grader checked the Salesforce notes but never looked at the DocuSign template. A submission that sent four contracts on the Standard template, instead of the required GDPR, HIPAA, SOC2 and Enterprise ones, passed 13 of 13 checks and scored 1.0. After the fix it scores 0.09.