Apple 提出 RISED:基于 rubric 的多环境智能体训练数据选择方法
RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
Apple 的论文,讲怎么在多个交互环境里同时训练一个 LLM 智能体,用 rubric 评分挑数据,全对全错的样本也能靠自蒸馏利用起来。
Apple 发布论文 RISED(Rubrics for Agentic Multi-Environment Selection and Self-Distillation),针对单一 LLM 智能体在多个交互环境中联合训练的问题。该方法用 rubric 评分比较不同环境下的 rollout 组,替代只依赖局部奖励信号的环境级数据选择策略。针对批次中同时出现的全失败和全成功 rollout 组缺少组间相对奖励信号的情况,RISED 引入自蒸馏机制利用这些数据。论文指出环境学习速率不同是现有 curriculum 方法失效的关键原因。
RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both…