AI模型精选

SG-TULA:非凸非光滑采样的驯化次梯度算法

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

精选理由

新采样算法SG-TULA能处理非光滑非凸问题,在GPT-2预训练上干赢了AdamW和Muon,还带理论保证。

AI 摘要

研究者提出次梯度驯化非调整朗之万算法(SG-TULA),用于处理势函数非光滑、超线性梯度增长且非凸的采样问题。该算法直接操作次梯度,无需计算昂贵的平滑步骤,并采用驯化技术保证显式格式稳定。在Wasserstein-2距离下推导了非渐近收敛界,所有常数随维度和逆温度显式给出,优于现有次梯度类朗之万算法。作者在GPT-2系列LLM的正则化预训练势上验证了假设,并展示SG-TULA的坐标提升变体在预训练中可与微调后的AdamW和Muon竞争。

原文 · arXiv cs.LG

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.