论文

研究证明带梯度裁剪的去中心化 SGD 在重尾噪声下可达到最优收敛率

Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

精选理由

搞分布式训练的可以看看这篇:它证明了裁剪过的去中心化 SGD 在重尾噪声下收敛率最优,还首次给出了线性加速证明,并解释了为什么裁剪比归一化更适合去中心化场景。

arXiv 论文研究了重尾噪声下去中心化优化的收敛问题,聚焦 clipped DSGD 方法。理论结果显示,在 p∈(1,2] 阶有界矩噪声条件下,clipped DSGD 在光滑非凸目标上以高概率和期望意义均达到 order-optimal 收敛率,并证明了随智能体数量的线性加速,这是首次针对带裁剪的去中心化方法给出该结论。论文还指出关键区别:normalized DSGD 可能不收敛,而裁剪保留了梯度幅度信息。技术核心是对 consensus gap 的精细分析,将网络效应归入高阶项。数值实验验证了理论结果。

原文 · arXiv cs.LG

Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

Heavy-tailed noise has been widely observed in modern machine learning, motivating the use of methods like gradient clipping and normalization. While these methods are well understood in centralized settings, much less is known in decentralized ones, where applying a nonlinearity to local gradients affects both optimization and consensus. Recent works on decentralized non-convex optimization have studied both clipping and normalization under heavy-tailed noise, with clipping yielding suboptimal rates and normalization needing local momentum or mini-batches to converge. This raises the question: can a baseline decentralized method using a nonlinearity achieve optimal convergence rates under heavy-tailed noise? We answer affirmatively with clipped decentralized SGD ($\mathtt{DSGD}$). For smooth non-convex costs under bounded $p$-th moment noise, $p \in (1,2]$, we show that clipped $\mathtt{DSGD}$ achieves order-optimal rates both with high probability and in expectation. Moreover, we establish a linear speed-up in the number of agents, which, to our knowledge, has not been shown for decentralized methods with clipping. The key technical ingredient is a sharp analysis of the consensus gap that exploits the structure of clipping, relegating network effects to higher-order terms. Our results highlight an important distinction between clipping and normalization in decentralized settings: while normalized $\mathtt{DSGD}$ can fail to converge, clipping retains magnitude information, enabling $\mathtt{DSGD}$ to be convergent and order-optimal. Numerical experiments validate our theory.