论文

MGFlow 提出统一分布训练框架,单步图像生成刷新 ImageNet 纪录

Unifying Distributional Training for One-Step Visual Generation

精选理由

单步生成图像的论文,用高斯混合做分布匹配,把 FLUX.2 4B 改成一步出图还比原来四步更好,搞生成的可以看看方法细节。

arXiv 论文提出把分布建模与匹配误差分离的统一理论框架,通过 Wasserstein 梯度流将全局目标与逐点特征更新联系起来,并由此推导出 MGFlow 方法。MGFlow 用高斯混合建模特征分布,同时支持最优传输和 score-based 两种匹配方式,通过质量约束的样本分配缓解模式坍塌。在 ImageNet 256×256 上,MGFlow 取得 1.45 FDr(pMF-H)和 1.64(JiT-H),超过 FD-Loss 基线。文本生成图像方面,MGFlow 将 FLUX.2 [klein] 4B 后训练为单步生成器,在 GenEval 和 PickScore 上超过原四步模型。

原文 · arXiv cs.LG

Unifying Distributional Training for One-Step Visual Generation

\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with \textbf{1.45} $\mathrm{FDr}^6$ on pMF-H and \textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/