论文

Probe-Space 预训练方法提升零阶训练效率

Probe-Space Preconditioning for Fast and Stable Zero-Order Training

精选理由

新方法1.5-SPSA让零阶训练比传统方法快1000倍,OPT-13B只需70步就超越MeZO 10万步效果。

研究人员提出1.5-SPSA方法,通过在probe-space中计算对角预处理器,显著提升零阶优化收敛速度。在Qwen3和OPT模型家族的6个数据集上,1.5-SPSA仅需70步即可超越MeZO方法10万步的性能,在SST-2基准上准确率提高3.1%。该方法结合8位打包随机生成器和分布式并行,实现在消费级GPU上稳定训练OPT-30B模型。

原文 · arXiv cs.LG

Probe-Space Preconditioning for Fast and Stable Zero-Order Training

Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this work, we evaluate two methods to close this gap. First, we show that reallocating training compute budget from many steps to large effective batch sizes with many perturbations (or probes) but fewer steps, allows 1SPSA (Spall, 1992) to outperform zero order methods like MeZO (Malladi et al., 2023) with less training compute. Next, we introduce 1.5-SPSA, adding a single "clean" forward-pass per step to 1SPSA to calculate a cheap diagonal preconditioner in probe-space, which improves convergence rate and convergence by down-weighting high curvature directions. Benchmarking on 6 post-training datasets on both Qwen3 and OPT model families, we show that 1.5-SPSA achieves State-of-the-Art results over previous ZOO solvers with much less optimization steps. For example, we train OPT-13B (for direct comparison to MeZO) and find 1.5-SPSA achieves +3.1% accuracy on SST-2 over both MeZO and BP in only 70 steps vs. MeZO's 100,000 steps. Finally, we combine an 8-bit-packing random generator, triton fused unpack/apply kernels, and distributed parallelism to achieve fast and stable training of models as large as OPT-30B in-place on commodity GPUs (e.g. A100).