别信那些基准排名——AISI发现给智能体多点token,表现就能飙升25%。新模型潜力更大。
英国AI安全研究所(AISI)在涵盖7项基准测试的研究中发现,标准评估通过限制计算预算系统性地低估了AI智能体的实际能力。在软件工程任务中,当token预算增加10倍时,成功率提升约25%。新模型受益最大,实际进展比之前测量结果陡峭约60%。
UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do
In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks, success rates jumped about 25 percent when the token budget was increased tenfold. Newer models benefit the most. Depending on the token budget, actual progress at the frontier is about 60 percent steeper than previous measurements suggested, according to AISI. The article UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do appeared first on The Decoder .
- @koltregaskes07-03 10:04原文