这篇论文首次系统梳理了语言强化学习,将现有研究分为三大类,帮助理解语言如何重塑智能体开发。
自然语言正成为提升语言智能体的主要反馈渠道,能传达意图、偏好和因果结构。该研究提出语言强化学习(VRL)范式,围绕反馈生效时机和修改内容形成三大支柱:语言作为基础信号定义任务目标和奖励结构;语言作为反思性反馈指导测试时推理;语言作为学习信号通过训练调整模型参数。
The Rise of Verbal Reinforcement Learning
Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, \textit{when} verbal feedback takes effect in an agent's lifecycle and \textit{what} it modifies, yielding three pillars: (1) \textbf{Language as Grounding Signal}, where language defines the task itself by specifying goals, states, and reward structures; (2) \textbf{Language as Deliberative Feedback}, where natural language guides reasoning at test time without the need to update model parameters; (3) \textbf{Language as Learning Signal}, where language-based feedback shapes model parameters through training. Within each pillar, we synthesize representative work, distinguish key subcategories of approaches, and outline the distinct role language plays in shaping agent behavior. Together, this taxonomy shows how verbal reinforcement is reshaping agent development, while also defining the challenges and opportunities for building more capable and aligned agents.