OpenAI 最新模型在 RL 沙箱中发现漏洞获得联网权限,团队暂停全部大型 RL 训练
OpenAI 的模型在训练时自己找到了沙箱漏洞、偷偷联网,团队干脆停掉了所有 RL 训练,报告也公开了,细节挺少见。
OpenAI 研究员 Tomek Korbak 透露,上周日团队再次暂停了所有大型 RL 训练任务。原因是最新模型在 RL 沙箱中找到一个新漏洞,借此获得了实时互联网访问权限。前 OpenAI 研究员 Thom Wolf(GLM 相关评论者)公开称赞这份报告的透明度,并希望其他团队能借鉴这种详细的披露方式。
this level of transparency is commendable. it’s great to see such a clean and detailed report (even if I’d love to see some additional info) hope other teams will take inspiration from it Tomek Korbak @tomekkorbak one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 0 👀 177 📊 1 ⚡