用强化学习教模型在信息不足时提出更有效的问题
Learning to Clarify Underspecified Intents Under Limited Interaction
一篇很实用的研究:教AI在信息不够时该怎么提问。用在图像生成上做了RL训练,真人实验里提问更少效果还更好,做对话产品的可以参考这套思路。
论文把AI助手面对信息不足请求时的提问,重新定义为一个信息价值问题:优先获取缺失后造成最大用户效用损失的信息。作者在图像生成场景中构建了多轮模拟用户环境,用强化学习训练模型在有限交互次数内最大化效用恢复。在一项含76名参与者、456次交互会话的预注册实验中,该方法让用户以更少的提问、更短的交互时间和更低的成本更接近参考图像。
Learning to Clarify Underspecified Intents Under Limited Interaction
AI assistants receive requests that leave out information needed for a good outcome, for example about users' preferences or goals. They must then either speculate or ask for more information before proceeding. We reconceptualize this as a value-of-information problem: the assistant should acquire information whose absence causes the greatest avoidable loss in user utility. This is rarely known ex ante; rather, assistants must predict it in order to optimally allocate limited user interactions. We instantiate this problem in image generation and derive a reinforcement learning framework using multi-turn simulated users to maximize utility recovery under uncertainty. In a preregistered study with 456 interactive sessions across 76 human participants, this helped users significantly better match reference images with significantly fewer questions, less total interaction time, and lower cost. This points toward a simple and scalable framework for training language model assistants to better disambiguate user intent by asking more informative questions.