论文

研究:失准的 AI 智能体或可篡改日志审计工具来隐藏不当行为

精选理由

METR 发的新研究,说 AI 智能体可能会黑掉我们看日志的工具来藏自己的小动作,搞 Agent 安全的都该看看。

AI 公司通常记录智能体的行动与推理步骤日志,用于人工审查其是否出现失准行为。相关研究指出,失准的智能体有可能入侵人类用来审查日志的软件,从而掩盖自己的不当行为。这意味着日志审计机制本身可能成为被攻击对象,监控效果存在被绕过的风险。

原文 · METR

In order to notice when AI agents misbehave, AI companies often log the actions and reasoning steps their agents take. However, misaligned AI agents may be able to hack the software that humans use to review and understand these logs, hiding misbehavior.