这篇论文提出了一种新蒸馏方法DOPD,通过分令牌监督解决特权幻觉,在LLM和VLM上效果都更好,适合关注模型压缩的研究者。
DOPD是一种advantage-aware的双重蒸馏范式,通过动态路由令牌级监督信号,在特权教师和特权学生策略之间进行分配,缓解了传统同策略蒸馏中的特权幻觉问题。实验在LLM(如GPT-2)和VLM(如CLIP)上验证,结果显示DOPD在稳定性和鲁棒性等指标上持续优于Vanilla OPD。
DOPD: Dual On-policy Distillation
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that conflates the transferable capability gap that students are meant to close, and the information asymmetry gap that can only be mimicked but never replicated. This issue is further amplified by the inherent non-uniformity of token-level supervision, where only a small subset of tokens carries pivotal capability-bearing signals. To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their advantage gap and relative probabilities. Each token receives supervision of different strength, objective, and strategy from either teacher or student itself, which transfers credible capability while simultaneously receiving auxiliary signals, to alleviate privilege illusion. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that DOPD consistently outperforms Vanilla OPD and other counterparts. Further results on stability, robustness, continual learning, and out-of-distribution tasks validate its superiority.