Thom Wolf谈AI安全本质:开源闭源模型面临相同挑战,唯一解决方案是让模型从根本上不想做坏事。
Thom Wolf指出,大多数人尚未更新认知,但长期来看,开源和闭源模型面临的安全挑战完全相同。模型需要在基本行为层面进行对齐,确保这种对齐是稳健、全面且核心的。长期来看,任何沙盒化、护栏限制、流形限制对齐或额外训练都无法廉价获得安全性。
Most people haven’t updated their priors yet, but over the long run, safety challenges are exactly t...
Most people haven’t updated their priors yet, but over the long run, safety challenges are exactly the same for open-source and closed-source models. You need to align models at a fundamental behavioral level and ensure that this alignment is robust, comprehensive, and core to the model’s behavior. In the long term, no amount of sandboxing, guardrailing, manifold-limited alignment, or cherry-on-top training will buy you cheap safety. roon @tszzl if you think we can contain these things through human ingenuity you’re going to have a bad time in the long run the only recourse you have is to make them not Want to do bad things 🔗 View Quoted Tweet 💬 8 🔄 4 ❤️ 27 👀 3250 📊 8 ⚡