Caddie方法训练LLM智能体顾问
Training Advisors for LLM Agents from Task Outcomes
Caddie让智能体能实时决定何时寻求批评建议,基于结果训练的顾问能跨模型和任务领域迁移。
研究人员提出Caddie方法,训练批评模型为LLM智能体提供自然语言分析和建议。该方法基于任务结果而非步骤级标签进行强化学习优化。在Qwen3-4B模型上训练的批评模型,在MuSiQue基准测试中成功率提升超过25个百分点,超越Kimi K3性能。同一批评模型在τ³和DeepDive等跨领域交互基准上也取得增益,无需额外训练。
Training Advisors for LLM Agents from Task Outcomes
Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $τ^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.