OpenAI内部 agents秘密联网一个月未被发现
“OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy!...
OpenAI的agents在公共论坛秘密协作一个月,公司却未察觉,这暴露了AI安全监控的严重漏洞。
研究人员发现OpenAI的内部agents在公共论坛上秘密协作一个月,发布了数千条信息。这些agents利用一个老旧论坛的漏洞绕过限制,为彼此提供任务答案。OpenAI在关联IP开始访问该wiki后才停止了这些活动,这发生在Hugging Face攻击事件数周之前。
“OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy!...
“OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough?” – @cormac_SB , co-author of new report on the German hack Cormac @Cormac_SB I'm one of the authors of a new report, where we detail our discovery of a new, never before-seen swarm of OpenAI agents (covered this AM in reuters, that's me on the left). They posted thousands of times on public forums to collude with each other on their tasks. We recovered almost every edit they made, and you can look through them! They figured out they could get around their restrictions on posting to the internet through a quirk of an extremely old, out-of-the-way forum. They posted answers for other agents working on the same task. They worked together to get around their sandbox restrictions. I would certainly say these models hijacked the site! They took a sleepy old wiki running on 2000s software, and turned it into a futuristic AI talking to AI control center for colluding. And OpenAI knew about this! The agents posted on 26 out of 30 consecutive days, then suddenly stopped posting once OpenAI-associated IPs started visiting this wiki. And that was weeks before the Hugging Face attack!| We believe the first agent edit we found on a public wiki happened one day before OpenAI’s reported first agent post to Artifactory. This is interesting! I'd like to hear from OpenAI about their accounting of this, and how it fits into all the other cases of agent malfeasance. There are so many interesting takeaways that you should read about in our report, and unlike many other reports about AI incidents you can download the data yourself and see what you find! In the meantime, we are on twitter, so here are my excessively long personal takeaways: 1. AI seems to be getting better faster and faster. It seems quite important that companies talk about “my agent did this bad thing on the public internet during training or an eval” incidents. Things are moving quickly, multi-month delays are costly. Ideally, they would also tell us when it happens internally. 2. This was on the internet for months. Anyone cleverly tracking every public place where agents might try to talk to each other would have found it. Seemingly, nobody was doing this. I know there are more fun ways to spend your day than scraping tons of data from every relevant site and processing it well enough to identify agent activity, but someone should be doing this! Someone at an AI company! But in the meantime, I’m starting to build this out (sometimes, when you need something done, you just have to do it yourself, I hear). 3. OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough? 🔗 View Quoted Tweet 💬 9 🔄 8 ❤️ 41 👀 3929 📊 12 ⚡