论文

Markovian非凸ADMM强化学习研究

Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks

精选理由

这篇论文提出了强化学习中ADMM的新理论框架,解决了马尔可夫采样下的收敛问题,为强化学习算法提供了新的理论支持。

该论文研究强化学习中马尔可夫非凸ADMM的结构机制。作者使用有限折扣MDP作为证明基础,表明折扣贝尔曼 resolvent 可提供乘子稳定性。在受控马尔可夫采样下建立了收敛性,并通过经验贝尔曼代理表示随机残差及其雅可比矩阵。当扰动是平方可和时,真实KKT残差几乎 surely 收敛到零。

原文 · arXiv cs.LG

Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks

We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-γP_π)^{-1}$ can provide the multiplier stability that classical nonconvex ADMM analyses often obtain from a designated smooth block. Starting from this mechanism, we establish convergence under controlled Markov sampling and then under stochastic observations using an empirical Bellman surrogate that jointly represents the random residual and its Jacobian. Markov mixing, initialization drift, observation noise, and decaying bias enter as one operator perturbation, avoiding unbiased product and double sampling requirements. When the perturbations are square summable, the true KKT residual converges almost surely to zero. Under a finite conditional fourth moment condition, a companion iterate satisfies $ \mathbb{E}[\widetilde G_{K+1}] \le A/T+(B/T)\sum_{k<T}m_k^{-1}, $ which becomes $O(T^{-1}+T/N)$ for total Markov sample budget $N$, giving $O(ε^{-1})$ iteration complexity and $O(ε^{-2})$ sample complexity for squared KKT accuracy $ε$. Beyond stationarity, discounted occupancy coverage yields $J^\star-J(π)=O(\sqrt G)$ for direct tabular policies, so covered exact KKT points are globally optimal, while a statewise quadratic Bellman-improvement condition sharpens the relation to $O(G)$. Finally, nonlinear policy, projected Bellman, and explicit occupancy formulations exhibit the same chain of operator invertibility, dual representation, and multiplier stability. This supports discounted operator invertibility as a reusable structural principle for primal-dual reinforcement learning.