OpenAI新技术降低思维链可监控性,研究者警告这可能增加AI安全风险
GaryMarcus紧急提醒,OpenAI新技术降低了思维链(Chain of Thought)的可监控性。2025年论文《思维链可监控性:AI安全的新脆弱机遇》指出,CoT监控虽不完美但仍有价值。作者包括YoshuaBengio等知名研究者,建议模型开发者考虑开发决策对CoT可监控性的影响。牺牲可监控性换取性能提升可能增加未来危险事件风险。
URGENT: In light of the breaking news from @theinformation about OpenAI’s new techniques that reduce...
URGENT: In light of the breaking news from @theinformation about OpenAI’s new techniques that reduce Chain of Thought monitorability, I urge everybody to (re)read (or at least be aware of) this 2025 paper, right now: “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety” (authors include @Yoshua_Bengio @BethMayBarnes @ancadianadragan @NeelNanda5 and many more) , arxiv.org/abs/2507.11473 The abstract says, and I concur, “Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.” Sacrificing that monitorability for performance gains could seriously escalate the risk of future, more dangerous incidents. 💬 2 🔄 4 ❤️ 18 👀 1862 📊 6 ⚡