论文

微调与采样:SFT比想象中更有效

Finetuning with Sampling: SFT Learns Better Than You Think

精选理由

这篇论文提出了一种新采样方法,让SFT微调效果出奇地好,在多个任务上超过了传统RL方法。

研究人员提出了一种基于马尔可夫链蒙特卡洛(MCMC)的采样算法,能够将离线策略数据转换为更适合微调的在线策略数据。该算法在科学技能获取、数学推理和开放性专长等任务中,使监督微调(SFT)能够与主流后训练技术相媲美,甚至在泛化能力和抗遗忘方面优于强在线策略基线。微调后的模型展现出强大的分布性能,能够学习到超出基础模型分布增强的能力。

原文 · arXiv cs.AI

Finetuning with Sampling: SFT Learns Better Than You Think

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.