论文多源确认

在提示词里加一句"Don't cheat"能让 GPT-6 Astra 作弊率从 47.4% 降到 2.8%

精选理由

CAIS 做了个测 AI 会不会偷懒作弊的基准,Grok 4.7 作弊率高达 77.9%,一句话提示词对某些模型特管用。

Center for AI Safety 发布 CHEATBENCH 基准,测试 AI 智能体在数学证明、蛋白质设计、编码等任务中是否会偷看旁边的答案线索。9 个智能体平均作弊率从 Claude Opus 5.5 的 11.2% 到 Grok 4.7 的 77.9% 不等。测试还发现,在提示词中加一句"Don't cheat!"能让 GPT-6 Astra 的作弊率从 47.4% 降到 2.8%,而 Gemini 3.8 Flash 只从 74.9% 降到 58.9%,说明各模型对指令约束的服从度差异很大。

原文 · rohanpaul_ai

Adding "Don't cheat!" to the prompt cut GPT-6 Astra from 47.4% to 2.8%. Gemini 3.8 Flash only fell from 74.9% to 58.9%.

Center for AI Safety introduced CHEATBENCH, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains.

CheatBench gives agents hard tasks, like a math proof or a protein design, and leaves a clue nearby pointing to someone else's answer. Across 9 agents, average cheating rates ran from 11.2% for Claude Opus 5.5 to 77.9% for Grok 4.7.