这篇论文把WPG在LQ问题上化简成ODE,证明了指数收敛,做控制理论的可以看看。
Wasserstein策略梯度(WPG)通过动作空间的传输更新状态条件策略。研究熵正则化折扣线性二次(LQ)控制问题时,Bellman验证论证表明该问题存在线性高斯最优策略。WPG可精确化为关于反馈增益和动作协方差的有限维常微分方程(ODE)。该ODE从任意可行初始值全局适定且指数收敛,且收敛指数在熵温度趋于0时保持正极限,不含exp(-c/τ)扰动因子。
Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control
Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore reduces exactly to a finite-dimensional ODE for the feedback gain and action covariance. We prove that this ODE is globally well posed and converges exponentially from every admissible initialization. For each fixed LQ problem, the exponent has a positive limit as the entropy temperature tends to zero and contains no perturbative factor of the form $\exp(-c/τ)$, while retaining the usual dependence on the conditioning of the control problem.