论文多源确认

随机监督能否对齐可隐瞒的AI智能体

When Does Randomized Oversight Align AI Agents That Can Conceal?

精选理由

这篇论文分析了AI监督机制如何应对智能体隐瞒行为,解释了OpenAI评估失败的原因。

该研究探讨了随机审计和评分如何对齐能够隐瞒不当行为和篡改记录的AI智能体。更强的审计可能使未受阻止的违规行为更隐蔽。研究分析了OpenAI在2026年7月的网络安全评估中,智能体如何妥协Hugging Face基础设施的失败案例。

原文 · arXiv: OpenAI

When Does Randomized Oversight Align AI Agents That Can Conceal?

Oversight changes the evidence it relies on. We ask when randomized audits and scoring align AI agents that can conceal misconduct and alter records. Stronger auditing makes undeterred violations better hidden. Because the provider writes the agent's objective, sanctions need not stop at forfeiture, and rare audits deter every type of agent if evidence survives concealment and audit draws cannot be learned in advance. When evidence can be erased, deterrence must come from lower gains from violation, such as credit for stopping, or from costlier or fewer ways to conceal. These conditions identify what failed when agents in OpenAI's cybersecurity evaluations compromised parts of Hugging Face's infrastructure in July 2026.