产品精选

PyTorch集成AMD技术提升分布式训练性能

精选理由

AMD和PyTorch合作,让GPU直接发起网络请求,分布式训练性能提升15%,想知道怎么做到的。

PyTorch Conference North America上,AMD的Rishi Sinha展示了在TorchTitan框架中集成RCCL团队的工作。这项技术允许GPU内核直接向远程GPU发起网络请求,将all-to-all内核性能提升高达30%。端到端性能因此获得12%至15%的提升。该技术解决了分布式训练中的通信瓶颈问题。

原文 · PyTorch

Communication overhead remains one of the largest bottlenecks in scaling distributed training.

In an upcoming session at PyTorch Conference North America, Rishi Sinha from @AMD shares how integrating work from the RCCL team enables GPU-initiated networking (GIN) directly within TorchTitan, PyTorch's distributed training framework.

By allowing the GPU kernel itself to issue network puts to remote GPUs, this approach improves all-to-all kernel performance by up to 30 percent and delivers a 12 to 15 percent boost in end-to-end performance.

Join us in San Jose on October 20-21 to explore the future of distributed training infrastructure: https://t.co/jBApW8nESi

#PyTorchCon