Salesforce AI Research 提出 Critical-State RL 训练多轮工具调用
Salesforce 这篇论文很有意思:训练多轮工具调用时只练关键那一步,在 BFCL v4 上能多拿 14 个点,做法讲得很清楚。
Salesforce AI Research 发布一篇关于强化学习训练多轮工具调用的论文。核心发现是只应训练那个真正改变结果的调用,而不是把奖励平摊到整条轨迹。Critical-State RL 用嵌套采样把当前动作引起的奖励变化与下游噪声分离,再用 contextual-bandit 更新只训练选中的那次调用。在 BFCL v4 missing-function 任务上,训练选中的那一轮带来约 14 点提升,而训练另一候选轮次则让准确率持平或下降。
Great paper from Salesforce AI Research on RL for multi-turn tool use.
The finding is that you want to train the one call where the action changes the outcome, instead of spreading reward across the whole trajectory.
(bookmark it)
When reward depends on later turns, much of its variation comes from what happens downstream.
Critical-State RL uses nested sampling to separate the reward variation caused by the current action from that noise, then trains only the selected call with contextual-bandit updates.
On BFCL v4 missing-function tasks, training the selected turn adds about 14 points. Training the other candidate turn leaves accuracy flat or lower.
Paper: https://t.co/vWtoTEVq1a