论文

ReBF模型解决强化学习速度自举难题

Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping

精选理由

ReBF解决了强化学习中的速度自举不可能三角问题,在多个基准测试中大幅提升性能。

Retimed Bellman Flows (ReBF)通过动态调整流时间查询教师评论家,解决了现有方法的结构性困境。该模型在合成马尔可夫决策过程中将Wasserstein-1距离降低至原来的7.7分之一。ReBF在38个OGBench和D4RL离线强化学习任务中表现优于现有流评论家方法。ReBF结合了重定时时钟和去耦合噪声生成,构建了条件无偏速度目标。

原文 · arXiv cs.LG

Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping

Flow critics learn return distributions by transporting Gaussian noise to Bellman endpoints via continuous velocity fields. While velocity bootstrapping stabilizes training by querying a successor teacher, existing methods face a structural dilemma: on straight paths, no residual-free same-time affine mapping can preserve Gaussian initial noise while maintaining an unbiased target. To overcome this limitation, we introduce Retimed Bellman Flows (ReBF). ReBF queries the teacher critic at a dynamically shifted earlier flow time, aligning intermediate student and teacher trajectories. By combining this retimed clock with fresh, decoupled noise generation, ReBF constructs a provably conditionally unbiased velocity target that preserves the Bellman fixed point and contracts under Wasserstein distances. Empirically, ReBF reduces $W_1$ distance to ground-truth return distributions by up to $7.7\times$ on synthetic MRPs and outperforms existing flow critics across 38 challenging OGBench and D4RL offline reinforcement learning tasks.