PDMD:一行代码改进视频扩散蒸馏,Wan2.1 上 4 步生成达 83.73 分
PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
一行代码就能修好 DMD 视频蒸馏的过饱和问题,Wan2.1 上 4 步出片还比原版高 1 分,做视频生成的可以直接看代码
DMD 蒸馏将视频扩散模型的去噪步数降到几步,但训练中 critic 误差会累积,导致画面逐步过饱和并出现伪影。论文提出 PDMD,通过把 DMD 更新中平行于 student-critic 端点残差的分量投影掉来过滤 critic 误差,并证明该残差是 critic 端点误差的无偏估计。PDMD 只需在 DMD 基础上改动一行代码,不增加额外损失、网络、数据或训练阶段。在 Wan2.1 上以 4 NFE 取得 VBench 总分 83.73,比同配置 DMD 高 1.03 分;在 MiniMax-H3 视频-音频联合生成上,VideoGen-Eval 视觉得分 83.17,6 项音频指标均优于对比的 4-NFE 模型。
PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.