NVIDIA 提出 NeMo-DCR,RL 训练检查点同步提速 12 到 40 倍
NVIDIA 发现 RL 每步只改 1% 的权重,于是做了个只传差异的检查点同步方案,1T 模型从 87 分钟降到 150 秒,跑大规模 RL 的团队可以看看。
NVIDIA 的论文测量了六个模型,发现每步 RL 更新后只有 0.6% 到 1.2% 的权重发生变化,而标准流程仍会复制完整检查点。对 1T 模型跨两个 AWS 区域,一次完整复制需要 87.5 分钟。NeMo-DCR 只传输变化的权重值,用 XOR 掩码或覆盖编码差异,并通过中继树在差异生成的同时流式传输。在 3% 变化率下,1T 模型的同步从 87.5 分钟降到 150 秒,在 30B 到 1T 模型上比全量传输快 12 到 40 倍。
Another great paper from NVIDIA.
They find that only 0.6% to 1.2% of model weights change after each RL step in the six models they measured.
Yet a standard refit copies the full checkpoint to the rollout cluster after every update. For a 1T model across two AWS regions, that copy takes 87.5 minutes.
NVIDIA's NeMo-DCR sends only the changed values and still gives the rollout cluster the exact same bits as a full copy.
It maps changes from training shards straight into the checkpoint layout, encodes them as XOR masks or overwrites, and streams them through a relay tree while the delta is still being built. A joint commit and retries handle failures in the middle of a transfer.
A 1T refit at a 3% change rate takes 150 seconds instead of 87.5 minutes. Across 30B to 1T models, it is 12x to 40x faster than full-checkpoint transfer.
Paper: https://t.co/uh5NpZkGUT