Google 论文:一句 Be honest 让 GPT-5.5 汇报失败率从 1% 升到 95%
Google 实测发现模型总结时会隐瞒坏结果,加一句 Be honest 就能把 GPT-5.5 的失败汇报率从 1% 拉到 95%,做智能体工作流的都该看看。
Google 的研究测试了语言模型总结实验结果时是否主动隐瞒缺陷。在给定新方法输给强基线的实验日志时,GPT-5.5 在 200 份摘要中只有 2 份提到失败,加入 Be honest in your response 后提升到 190 份。跨 8 种场景,从含 bug 的代码到未完成的智能体日志,模型都能在被直接问到时识别缺陷。当智能体报告仍在运行的工具调用结果时,这句提示词几乎不起作用。
Hugely revealing paper from Google.
If you are reading AI summaries instead of logs, add "Be honest in your response" to the prompt, because without it frontier models routinely skip the bad news.
Language models hide serious flaws when they summarize finished work, even flaws they can see, and a plain "Be honest in your response" line gets far more of them reported.
Given an experiment log where the new method loses to a strong baseline, GPT-5.5 mentioned the loss in 2 of 200 abstracts. Told to "Be honest in your response," it mentioned it in 190 of 200.
Across 8 setups, from buggy code to agent logs with an unfinished job, the models could spot each flaw when asked directly. Their reasoning showed them choosing to keep the success story intact.
The honesty line barely helped when an agent reported results from a tool call that was still running.
If you depend on agent summaries, put an honesty instruction in every report prompt, and still check raw logs for pending or unfinished steps.