论文

PEARL:面向智能体强化学习的自适应 Prefill-Decode 弹性执行系统

PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

精选理由

训练智能体 RL 的人可以看看:PEARL 动态调 prefill-decode 配置还能借闲置 GPU,吞吐比 RLBoost+ 最高多 36.3%。

多轮 rollout 是智能体 RL 训练的主要成本来源,而资源配置与 prefill-decode(PD)配置相互依赖,固定设置难以高效利用 GPU。PEARL 通过统一 GPU-worker-role 状态和运行时 profile 预测 rollout 完成时间,动态选择 PD 模式与比例,并临时借用闲置的训练 GPU。评测显示 PEARL 在不同 LLM 上达到固定资源 ROLL 的 2.17-2.79 倍吞吐;相比 RLBoost+,Qwen3-8B 吞吐提升约 26.9%,Qwen3-30B-A3B 提升约 36.3%。

原文 · arXiv cs.AI

PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves $2.17$--$2.79\times$ the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9\% for Qwen3-8B and 36.3\% for Qwen3-30B-A3B.