IDRF:掩码离散扩散模型的逆蒸馏奖励微调框架
IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
一篇扩散模型微调的论文,教你怎么给少步掩码扩散模型做奖励微调,还能少 32 倍去噪步数,做生成方向的可以看看。
arXiv 论文提出 IDRF,用于对少步掩码离散扩散生成器做奖励微调。方法上用逆蒸馏正则替代不可解的序列级 KL 惩罚,并证明在最优辅助去噪器下该损失可上界序列 KL 散度。IDRF 将少步生成建模为有限时域 MDP,采用裁剪策略梯度优化奖励,无需参考模型 rollout。在 DNA、图像和文本生成任务上,相比参考模型最多减少 32 倍去噪步数,同时缓解 reward hacking 并保持样本质量。
IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.