论文精选

Muon优化器的哈密顿概率梯度流视角

Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer

精选理由

这篇论文为Muon优化器提供了严格的数学基础,揭示了其与哈密顿动力学的深层联系。对优化理论研究者或想深入理解Muon工作机制的深度学习从业者,值得细读。

AI 摘要

本文从概率梯度流的角度重新审视了Muon优化器,将其视为一种正则化的镜像/近端步骤。作者发现正则化的正交化映射是核范数的光滑Fenchel对偶平滑的梯度,从而将Muon更新与动量作为对偶坐标联系起来。通过将Muon从单矩阵参数提升到有限粒子概率目标,推导出惯性连续时间极限,并建立了相空间平均场方程。该流被证明是一种阻尼哈密顿概率动力学,其哈密顿能量单调递减。在额外假设下,论文证明了目标间隙的指数收敛速率,并研究了平均场极限方程的适定性和传播混沌保证。最后,将公式扩展到希尔伯特值特征映射,得到适用于平滑Transformer混合专家模型的块状Muon概率流。

原文 · arXiv cs.LG

Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer

We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dual smoothing of the nuclear norm. This identifies the (regularized) Muon update as a mirror/prox step in the update variable, with momentum acting as the dual coordinate. We use this structure to lift Muon from a single matrix parameter to finite-particle probability objectives of the form $J(ρ)=R\left(\int F d ρ\right)$, a setting motivated by mean-field descriptions of neural-network training, and derive the inertial continuous-time limit. Using this structure, we derive the finite-particle continuous-time limit under the inertial scaling of step size and momentum, and then pass to a phase-space mean-field equation over probability laws on parameter-momentum pairs. The resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically. While the target objective itself need not be monotone along the inertial Muon dynamics, under additional gradient-dominance, bounded-momentum, and curvature/alignment assumptions, we obtain continuous and discrete-time exponential convergence rates for the objective gap. We also study the well-posedness of the mean-field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.