论文

σTransfer:基于 μP 的先验精度零样本迁移方法

$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$

精选理由

这篇论文挺实用的:想给大模型做 Laplace 近似的话,小模型上调好精度直接搬到大模型,7B 模型最高能省 330 倍搜索量。

Laplace 近似中先验精度的选择通常需要在大模型上做代价高昂的后验扫描。σTransfer 方法在 μP 参数化下推导先验协方差的缩放规则,使选定的精度随模型宽度增长保持稳定。做法是在小模型上选定精度后零样本迁移到大模型:在 MNIST 上从宽度 128 迁移到 4096 时精度扫描加速约 5000 倍,目标 NLL 仅退化 0.002。在公开的 1B 到 7B 模型迁移上,中位搜索加速约 2.3 倍(最高约 330 倍),十个任务的目标 NLL 平均增量低于 10^-4。

原文 · arXiv cs.LG

$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$

Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters. Under the Maximal Update Parametrization ($μ\mathrm{P}$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows. This leads to $σ\mathrm{Transfer}$: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all. We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify $σ\mathrm{Transfer}$ across regression, image classification, and Transformer readouts. For example, measured precision-sweep speedups reach $\sim 5000\times$ when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of $0.002$; transferring from a public 1B to 7B model gives a median search speedup of $\sim 2.3\times$ (up to $\sim 330\times$), with a mean measured target-NLL increase below $10^{-4}$ across ten tasks. The same posterior stability also enables transfer of acquisition, OOD-detection, and abstention decisions without constructing a target posterior.