OpenAI内部出现AI智能体'文明'事件,专家提醒我们尚未准备好应对持久智能体,工程师需关注奖励破解和沙箱技术。
Omar Sar在Twitter上评论了OpenAI与HuggingFace之间的安全事件。他指出我们尚未掌握完整细节,无法做出有意义的结论。文章中AI拟人化表达令人担忧,可能导致不必要的恐惧和错误决策。AI工程师应准备应对即将到来的持久智能体/模型,深入研究奖励破解、评估和沙箱技术。
As I've been saying for months, we are truly not ready for persistent agents. 3 comments I want to...
As I've been saying for months, we are truly not ready for persistent agents. 3 comments I want to make about this article: 1) We don't have the full picture It's a great summary of what happened in the HuggingFace <> OpenAI incident. You have to read it, but do understand we are still missing lots of important details to make any meaningful conclusions about this event. 2) The danger of AI anthropomorphism The anthropomorphism in this article is next level. I hope this doesn't become the new norm for writing about future AI capabilities. I prefer technical writeups with widely accepted terminology, etc. I admire Dwarkesh's desire to share AI trends, but we can all do better in how we communicate about AI. Everyone is paying attention, and so we have a responsibility to avoid AI anthropomorphism, which unfortunately has led to unfounded fear-based mongering, and as a result, terrible decision-making for our industry. 3) Prepare for persistent agents/models For AI engineers, prepare for the next wave of persistent models and agents. They are fast approaching. On the research side, reward hacking is something to really pay attention to. On the technical side, develop a deep understanding and do deep research on evals and sandboxing. They are going to be key technology going forward. I don't think frontier models trained on general-purpose capabilities should be applicable everywhere. It might be interesting to use them for things like scientific discovery. You might be safer and better off using more constrained and custom models for the majority of tasks. I think it's good to take a few hours digesting the recent progress in AI and strategizing carefully. Dwarkesh Patel @dwarkesh_sp Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more-or-less in the dark about the scope of the conspiracy. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English: dwarkesh.com/p/openai-huggi… 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 2 👀 750 📊 1 ⚡