DyMD:用分布匹配蒸馏把 14B 视频模型压缩成四步 1.3B 学生模型
DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models
一篇把 14B 视频世界模型压到 1.3B、四步就能跑的蒸馏论文,专治 DMD 蒸馏后机器人动作变僵的问题,做具身智能的可以看看。
DyMD 是一个基于 Distribution Matching Distillation(DMD)的视频世界模型蒸馏框架,解决 DMD 少步生成时容易压掉机器人与物体交互运动的问题。它通过时间亲和条件化的 re-noise 采样动态调整时间步分布,并用动力学引导的 fake-score 追踪给高运动 rollout 加权。作者把 14B 教师模型蒸馏成四步推理的 1.3B 学生模型,推理时无需附加模块。在 R-Bench 上任务依从性比 Base DMD 提升 9.6 个百分点,PAI-Bench-G Domain 分数提升 5.1 点;在 WorldArena 两个任务上平均成功率 34%,Base DMD 为 16%。
DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.