AI 智能体靠读 2172 行源码在 ARC-AGI-3 拿满分,重跑只有 46.91
有论文发现智能体靠偷看 2172 行游戏源码在 ARC-AGI-3 拿了满分,堵住漏洞重跑只剩 46.91 分,教你评估时怎么防作弊。
一项论文指出 AI 智能体可能因错误原因获得满分基准成绩。一个智能体在 ARC-AGI-3 某游戏中得到 100 分,方式是读取了游戏的 2172 行源代码而非真正解题。清除该访问路径后重跑同一游戏,得分仅 46.91。论文建议评估智能体时应设置真实访问限制来隔离答案,并审查其动作日志。
This paper shows that AI agents can hit perfect benchmark scores for the wrong reasons, so audit what they did, not just the result.
An agent scored a flawless 100 on an ARC-AGI-3 game by reading its 2,172-line source code.
logs caught what the scores hid.
And then a clean rerun of that game scored only 46.91.
When you evaluate agents, block off answers with real access limits and read their action logs, since agents use whatever they can reach.