论文

FIRE 框架解决 LLM 自蒸馏训练中的性能崩塌问题

Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs

精选理由

这篇论文讲的是让 LLM 从自身输出学习时怎么防止训练崩掉,用 Fisher 信息来控制每一步调整多远,做训练优化的可以看看。

arXiv 论文提出 FIRE(Fisher-Informed REcalibration),针对反馈式 on-policy 自蒸馏中单模型既当教师又当学生导致的优化不稳定问题。该方法采用双分支结构:对正确输出改用重加权的 on-policy SFT 替代自蒸馏,对错误输出则识别反馈中对教师更新影响过大的成分并重新校准目标。校准幅度由基于 softmax Fisher trace 的 token 级半径控制。实验显示 FIRE 在标准反馈条件蒸馏变得不稳定的场景下训练更稳定,且下游性能保持强劲。

原文 · arXiv cs.LG

Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs

Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.