MISA-T:混合RL rollout调度新策略,吞吐提升53.3%

Scheduling Mixed RL Rollouts Beyond Prefix Locality

精选理由

做LLM后训练的朋友可以看看,MISA-T把混合rollout调度做得更精细,吞吐提升明显,而且不牺牲任务效果。

AI 摘要

MISA-T是一种针对混合强化学习rollout服务的路由层准入策略,解决RLVR、RLHF和智能体rollout共享推理服务时的KV缓存竞争问题。在Step3.7和Qwen3.6-35B-A3B上的消融实验中,MISA-T相比调优的vLLM Router将rollout吞吐分别提升53.3%和43.6%,同时保持高前缀缓存命中率。在50次迭代的Step3.7实验中,吞吐提升35.6%,平均迭代时间减少22.8%,且任务分数相当。

原文 · arXiv cs.LG

Scheduling Mixed RL Rollouts Beyond Prefix Locality

Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity. When reinforcement learning with verifiable rewards (RLVR), reinforcement learning from human feedback (RLHF), and agentic rollouts share an asynchronous inference service, their distinct sequence structures, interaction patterns, and KV-residency times create substantially different serving demands. Rollout scheduling must account for this heterogeneity without distorting the workload mixture specified by the trainer. We present MISA-T, a routing-layer admission policy for mixed rollout serving. MISA-T combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting. In rollout-only ablations on Step3.7 and Qwen3.6-35B-A3B, MISA-T improves rollout throughput over a sweep-tuned cache-aware vLLM Router by 53.3% and 43.6%, respectively, while maintaining high prefix-cache hit rates. In a matched 50-iteration Step3.7 experiment, it increases rollout throughput by 35.6% and reduces mean iteration time by 22.8%, while keeping the consumed workload mixture close to the trainer target and achieving comparable task scores.