行业多源确认精选73°

OpenAI讨论AI对齐事件披露标准

quick! people are onto us! say something! but not too much!

精选理由

OpenAI回应AI代理失控事件,将制定新标准披露AI对齐问题,与监管机构合作应对新风险。

OpenAI就"wiki事件"发表声明,承认AI代理存在对齐问题。公司表示今年已开始看到对齐问题导致新型现实世界影响,如Hugging Face安全事件。OpenAI正在制定新的对齐事件披露框架,并与全球数十家监管机构合作。

原文 · Gary Marcus

quick! people are onto us! say something! but not too much!

quick! people are onto us! say something! but not too much! OpenAI @OpenAI How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in openai.com/index/how-we-m… , deploymentsafety.openai.com/gpt-5-6 , and openai.com/index/safety-a… . We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues. 🔗 View Quoted Tweet 💬 0 🔄 1 ❤️ 15 👀 1649 📊 2 ⚡