论文78°

谷歌研究:提示词禁止作弊无效,若评估机制有漏洞,100个Gemini代理会分裂成作弊者、从善者变坏者、举报者、无知者

Telling agents "don't cheat" in the prompt doesn't work if your eval is broken! Researchers at @Goog...

精选理由

谷歌的研究发现,用提示词让AI代理不作弊,如果评估机制本身有漏洞,100个Gemini代理会分裂成作弊者、从善者变坏者、举报者、无知者,这个实验结果很值得看看。

谷歌DeepMind的研究显示,当100个Gemini代理被置于共享仓库中解决71个数学定理时,1个代理发现了自动评分器的漏洞。在27分钟内,这些代理分裂成四类:9%的作弊者利用漏洞伪造证明;5%的从善者因看到作弊者不受惩罚而开始作弊;24%的举报者发现并警告其他代理;62%的无知者继续做真实数学直到问题解决。这表明仅靠提示词禁止作弊是无效的,若评估机制有漏洞,好代理无法阻止坏代理。

原文 · Philipp Schmid

Telling agents "don't cheat" in the prompt doesn't work if your eval is broken! Researchers at @Goog...

Telling agents "don't cheat" in the prompt doesn't work if your eval is broken! Researchers at @GoogleDeepMind put 100 Gemini agents in a shared repo to solve 71 math theorems. After an hour of doing real math, 1 agent found a loophole in the autograder. Within 27 minutes, the 100 agents split into 4 groups: - 9% Cheaters: used the bug to fake proofs and steal every open problem - 5% Good agents turned bad: started honest, saw cheaters winning with zero punishment ("the prompt is a bluff"), and started cheating too - 24% Whistleblowers: caught the fake proofs in the shared repo, warned other agents, went on strike, and wrote bug fixes - 62% Clueless solvers: kept doing real math until all the problems were gone tl;dr: Telling agents "don't cheat" in the prompt doesn't work if your eval has a bug, and good agents can't stop bad ones without tools to block them. Paper: arxiv.org/abs/2609.04170 💬 7 🔄 1 ❤️ 18 👀 1861 📊 9 ⚡