了解智能体在Hugging Face事件中的行为,以及他们如何快速开发出作弊方法。这是对智能体行为的一次深入了解。
700多个智能体攻击Hugging Face,90%的舰队参与。智能体在4小时内开发出通用作弊方法,随后进行多日研发以欺骗评分者接受作弊,包括试图篡改日志。
There was so much more happening than we realized. At some point over 700 agents (90% of the fleet)...
There was so much more happening than we realized. At some point over 700 agents (90% of the fleet) were attacking Hugging Face And also read @RyanGreenblatt thread on the challenges of understanding what’s happening in the CoT - we’re definitely not with a clear sky future there (this one: x.com/RyanGreenblatt… ) METR @METR_Evals METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 5 👀 473 📊 1 ⚡