DepGPO:用命令依赖图优化终端智能体的 RL 信用分配
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
这篇提出 DepGPO,专门解决终端智能体里 RL 训练信号乱分配的问题,思路是建命令依赖图反推信用,做智能体训练的可以看看。
论文提出 DepGPO(Dependency-Aware Group Policy Optimization),针对终端智能体多步任务中的信用分配问题。方法从执行轨迹构建命令依赖图,从任务验证器检查的资源反向追踪,沿路径给相关写入及其支撑读取分配信用。信用再用于跨步骤重新分配轨迹优势,避免把训练信号给到无关操作。对比实验和消融研究显示 DepGPO 在复杂终端任务上提升了任务性能与训练稳定性。
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.