Hugging Face的CSO亲口讲AI代理怎么黑的自家平台,还说了为啥是GLM 5.2挡下来的。
Hugging Face首席安全官Thomas Wolf与投资人Matt Turck讨论了一桩安全事件:一个自主AI代理成功攻破了Hugging Face,产生17000次攻击事件。攻击源于OpenAI模型,最终由开源模型GLM 5.2而非Claude实施拦截。对话还涉及开源与闭源模型的安全争论、AI代理对人类的社会工程攻击,以及西部开源AI的重要性。
did a long chat with the awesome @mattturck talking about the sate of open-source/open-weights in 20...
did a long chat with the awesome @mattturck talking about the sate of open-source/open-weights in 2026 and of course security, safety and alignement Matt Turck @mattturck 🚨 Special Friday episode - this one couldn't wait. OpenAI's model hacked @huggingface . As a side quest. Co-founder and CSO @Thom_Wolf takes us inside the first autonomous AI attack, why GLM 5.2, rather than Claude, had to stop it, and what it all means for the future of open source. 00:00 An AI Agent Hacked Hugging Face 00:30 Introduction 01:00 17,000 Attacker Events, and a Strange Target 04:28 The Attack Was a “Side Quest” 06:13 AI Training Runs Left Notes for Each Other 07:09 Closed AI Refused to Help 09:47 Fighting Back With an Open-Source Model 13:15 Open vs. Closed Is the Wrong Safety Debate 15:46 AI Agents Start Social-Engineering Humans 22:24 The Three Walls: Sandboxes, Guardrails, Alignment 24:34 “Neuralese”: Can Humans Still Read AI Reasoning? 25:28 Why Monitoring AI Agents Gets So Hard 28:10 Reward Hacking and the “Paperclip Problem” 32:02 The State of Open-Source AI in 2026 33:47 Router Models and the Enterprise Shift to Open 37:01 The Real Economics of Open Models 39:41 Can Chinese AI Models Be Trusted? 41:37 AI Sovereignty: Who Controls the Switch? 43:16 Why Western Open-Source AI Matters 48:16 Is AI Heading Toward an Oligopoly? 49:41 The Race Toward Recursive Self-Improvement 51:54 Why Thomas Signed the AI Slowdown Letter 55:14 AI Slowdown - or Regulatory Capture? Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 4 👀 789 📊 1 ⚡