这篇把GFlowNets训练用信息几何重新梳理了,用Fisher-Rao度量做自然梯度,还给了三种能实际算的场景,图模型工具也用上了,搞生成流网络的可以看。
生成流网络(GFlowNets)是面向离散及混合离散-连续对象的摊销推断框架,其训练仅需一个未归一化的奖励函数。本文将GFlowNets的前向策略视为轨迹采样器,证明其内在一阶几何由Fisher-Rao度量给出,并采用自然梯度作为局部更新方向。论文推导出轨迹Fisher信息矩阵的逐条件二阶矩分解,揭示时间分数交互何时消失、共享参数化下何时产生稠密耦合。据此划分出三种计算模式:精确Fisher信息可解、可用蒙特卡洛估计、以及可利用目标局部性或分解结构的情形。对第三种情形,论文用图模型工具(如精确边缘化、分隔符方法、置信传播)近似自然梯度更新,将目标结构转化为优化几何。
Information-Geometric Forward Policy Training in GFlowNets
Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward. In this work, we formulate forward-policy training in GFlowNets through the information geometry of the induced trajectory sampler. Treating the forward policy as an induced trajectory sampler, we show that its intrinsic first-order geometry is given by the Fisher-Rao metric of the trajectory family, and that the associated natural gradient provides the canonical local update whenever the corresponding Fisher information is computable or accurately approximable. We derive an exact decomposition of the trajectory Fisher into per-step conditional second moments, which clarifies when temporal score interactions vanish and when dense couplings remain under shared parameterisation. This leads to three computational regimes: settings with tractable exact Fisher information, settings where Monte Carlo estimators of the expected Fisher are sufficient, and structure-exploitable settings in which target locality or factorisation yields accurate approximations of the Fisher expectation. In the latter case, graphical-model tools such as exact marginalisation, separator methods, and belief propagation provide principled surrogates for natural-gradient updates. The resulting framework turns target structure into optimisation geometry and yields a tractable route to structure-aware forward-policy training in GFlowNets. We illustrate the framework empirically through examples comparing convergence and exploration behaviour under Riemannian and Euclidean optimisation.