CoreWeave RL Rollouts 实现训练中热加载权重,BrowseComp 提升 8.5 个点
做 RL 训练的朋友看下,CoreWeave 这个功能把换检查点的时间砍掉约 15 倍,Nemotron 实测 BrowseComp 涨了 8 个多点。
CoreWeave 推出 RL Rollouts,在强化学习训练循环中把新检查点权重热加载进线上部署,不中断进行中的请求,速度约为重新部署流程的 15 倍。NVIDIA 和 You.com 用它对 Nemotron 3.5 Lightning 做搜索任务后训练,BrowseComp 准确率从 36.97% 提到 45.45%,工具调用次数减少 30.24%。此前每次更新检查点都要重新部署,训练器只能等待。
In reinforcement learning, inference is part of the training loop. Every checkpoint used to mean a redeploy, and the trainer waited.
CoreWeave RL Rollouts load the new weights into a live deployment without touching in flight requests, about 15x faster than a redeploy cycle.
@nvidia and @youdotcom used it to post train Nemotron 3.5 Lightning for web search and lifted BrowseComp accuracy from 36.97% to 45.45% while cutting tool calls by 30.24%.
Full breakdown here: https://t.co/cTcqSRGeCB