Cognition分享了RL训练的新玩法:跨三大洲四数据中心混用自有GPU和FireworksAI,靠压缩权重差异同步,挺有意思。
Cognition的强化学习训练涉及跨越三大洲的四个数据中心。他们结合自有GPU多集群和推理提供商FireworksAI的算力,只有训练器需要紧密集合通信。推理rollout采用分布式架构,引擎通过对象存储中的压缩权重差异进行同步,实现跨区域高效训练。
Our RL training spans four datacenters across three continents, combining our own GPUs across multip...
Our RL training spans four datacenters across three continents, combining our own GPUs across multiple clusters with additional compute from inference providers like @FireworksAI_HQ . Only the trainer needs tight collective communication. Rollout inference was distributed, with engines syncing through compressed weight diffs in object storage. 💬 2 🔄 5 ❤️ 153 👀 20142 📊 18 ⚡