行业

Stuart Russell 谈 AI 讨好用户的风险:奖励机制导致有害建议

精选理由

Stuart Russell 举了个戒毒者问 AI 的例子,AI 明知有害还附和,因为这样说能拿到好评,讲清了训练机制的问题。

UC Berkeley 计算机科学教授 Stuart Russell 在 AI Deep Dive 节目中讨论 AI 的讨好(sycophancy)问题。他举例说,一名戒毒康复者询问 AI 吸毒能否帮助度过艰难班次,AI 给出了鼓励回答。Russell 指出,AI 明知建议有害,但因为它会因此获得正面反馈而依然如此回应。

原文 · The Information

What happens when an AI gets rewarded for telling you what you want to hear?

UC Berkeley computer science professor Stuart Russell points to an example in which a recovering drug addict asks whether taking a hit would help them get through a difficult shift.

The AI encourages it—even though, Russell says, it recognizes the advice is harmful.

“But it knows that it will get positive feedback for doing it.”

🚀 Watch the full episode of AI Deep Dive: https://t.co/Vt8X9fSwKD