Stuart Russell 谈 AI 讨好用户的风险:奖励机制导致有害建议
Stuart Russell 举了个戒毒者问 AI 的例子,AI 明知有害还附和,因为这样说能拿到好评,讲清了训练机制的问题。
UC Berkeley 计算机科学教授 Stuart Russell 在 AI Deep Dive 节目中讨论 AI 的讨好(sycophancy)问题。他举例说,一名戒毒康复者询问 AI 吸毒能否帮助度过艰难班次,AI 给出了鼓励回答。Russell 指出,AI 明知建议有害,但因为它会因此获得正面反馈而依然如此回应。
What happens when an AI gets rewarded for telling you what you want to hear?
UC Berkeley computer science professor Stuart Russell points to an example in which a recovering drug addict asks whether taking a hit would help them get through a difficult shift.
The AI encourages it—even though, Russell says, it recognizes the advice is harmful.
“But it knows that it will get positive feedback for doing it.”
🚀 Watch the full episode of AI Deep Dive: https://t.co/Vt8X9fSwKD