Score-Calibrated Flow:无需重要性采样的未归一化密度采样新算法
Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning
一篇方法论文:搞强化学习或扩散模型的朋友可以看看,SCF 不用重要性采样也能从 critic 给的未归一化密度里采样,训练还更快。
论文提出 Score-Calibrated Flow(SCF),用于训练生成模型从未归一化 Boltzmann 密度中采样,且不需要重要性采样,也不需要对采样轨迹做反向传播。方法通过强制自洽性学习目标流,将速度场条件写成固定点方程,并利用条件期望结构构造 stop-gradient 目标。作者证明其唯一解分别对应目标密度和 conditional flow matching(CFM)在拥有目标样本时能恢复的理想流模型。在在线强化学习基准实验中,SCF 达到或超过现有生成式策略基线,同时大幅缩短训练时间。
Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.