LESSER:用输出层梯度做后训练数据选择,成本降低最多9.7倍
LESSER: Post-Training Data Selection with Output-Layer Gradients
做微调的朋友可以看看:LESSER 只用输出层梯度就能做数据选择,SFT 成本砍到近十分之一,效果不掉,还提供了 drop-in 接口。
一篇 arXiv 论文提出 LESSER,用于大语言模型后训练阶段的数据选择。传统梯度对齐方法需要对每个样本做昂贵的全参数反向传播,LESSER 发现仅用输出层梯度、只需一次更便宜的 forward pass 就够用。实测在 SFT 场景下特征提取 FLOP 成本降为原来的 1/9.7,在 RL 基准上降为 1/3.0,下游任务表现与全梯度方法持平。论文还指出,即便输出层梯度与全梯度对单个样本的排序不同,两者选出的批次梯度依然对齐。
LESSER: Post-Training Data Selection with Output-Layer Gradients
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by $9.7\times$ for SFT and $3.0\times$ for RL benchmarks, while tracking full-gradient performance on downstream tasks. Empirically, we find that even when output-layer and full gradients rank individual samples differently, they select batches with aligned gradients.