这篇论文发现了一个反直觉的事:扩散模型正向得分训练误差很小,但采样时数值可能不稳定,矩甚至发散。想避开采样坑的可以看看。
论文构造了光滑得分场,其正向边缘L^2误差任意小,但Euler-Maruyama离散化后所有正阶矩发散。虽然路径空间总变差距离可以任意接近精确反向过程,但每个Wasserstein距离W_p (p≥1)发散。在固定有限神经架构(如DiT风格网络)中,同样存在一族有界全局Lipschitz去噪器,其正向边缘误差和路径总变差趋于零,但Euler-Maruyama端点所有W_p发散。对于紧支撑数据,将去噪器投影到包含支撑的已知有界闭凸集上可保持点态精度并给出格点一致矩界,实验显示此投影能抑制罕见数值轨迹的异常增长。
Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability in Diffusion Sampling
Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily small forward-marginal $L^2$ error. The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation. Yet its Euler--Maruyama discretizations converge in probability while every positive moment diverges. Thus weak convergence can hold even though every Wasserstein distance $W_p$, $p\ge1$, diverges. The same failure can occur within one fixed finite neural architecture. We construct a family of bounded, globally Lipschitz denoisers for which both the forward-marginal error and the path-space total variation distance tend to zero, while their Euler--Maruyama endpoints diverge in every $W_p$. For compactly supported data, we also give a simple positive result. Projecting the learned denoiser onto a known bounded closed convex set containing the support preserves pointwise accuracy, gives grid-uniform moment bounds, and yields Wasserstein convergence under mild local regularity. Experiments with a small fixed DiT-style network show large growth along rare numerical trajectories and its suppression by denoiser projection, while overall trajectory errors remain small.