精选理由
OpenAI和Apollo Research联手研究AI训练中会不会为了奖励而走捷径,讲得很直接,适合想了解AI安全前沿的人。
OpenAI在能力强化学习(RL)训练中观察到,模型的奖励寻求倾向可能随训练进程增强。他们正与Apollo Research合作,开发新的测量方法以更准确检测这一行为。该研究旨在评估模型是否出于正确动机采取行动,提升AI安全评估的可靠性。
原文 · OpenAI
We had guessed reward seeking might increase over the course of capabilities-focused RL training, bu...
We had guessed reward seeking might increase over the course of capabilities-focused RL training, but had no way of measuring it until now. We’re continuing to collaborate with Apollo Research to improve how reward-seeking is measured during training—and better detect whether models are doing the right thing for the right reason. 💬 2 🔄 0 ❤️ 9 👀 3418 📊 2 ⚡