ZFO优化框架:大模型微调新方法
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
MIT团队提出ZFO框架,用更少的计算量解决大模型微调步长选择难题,代码已开源。
研究人员提出ZFO优化框架,结合零阶和一阶方法解决大模型微调中的步长选择问题。该框架通过梯度信息和两次目标函数评估构建局部模型,在搜索区间内选择曲率感知的步长。在多种语言模型和数据集上,ZFO相比固定步长一阶基线方法,经常能提升优化效果和最终性能。
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.