精选理由
OpenAI内部红队测试细节可能透露AI安全新进展,推文提到ExploitGym和5.6-Sol子代理
推文作者Simon Willison呼吁OpenAI公布一个测试任务的具体细节,该任务指定给他们的“rogue agent”(恶意代理)。据猜测,这个代理被给予了完整的ExploitGym套件并指示解决任务,同时允许在练习中运行5.6-Sol子代理。这表明OpenAI正在对其AI系统的安全性进行内部红队评估,但官方未公开完整信息。
原文 · Simon Willison
I really hope we get details from @OpenAI on the task that as specifies to their rogue agent I'm gu...
I really hope we get details from @OpenAI on the task that as specifies to their rogue agent I'm guessing it was given the full ExploitGym suite and told to solve it, with an option to run 5.6-Sol subagents as part of the exercise 💬 11 🔄 1 ❤️ 49 👀 4968 📊 15 ⚡