论文精选

分位数时序差分学习有限样本分析

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

精选理由

强化学习新论文,分析QTD算法的有限样本性能,揭示了波动与样本复杂度的关键区别。

AI 摘要

该研究建立了表格分布强化学习中同步分位数时序差分学习的全局有限样本保证。研究证明步长α_t=c(t+1)^{-a}(a∈(1/2,1))下,最终迭代波动阶为T^{-a/2}/√(1-γ),与分位数数量无多项式依赖。研究区分了局部随机波动和全局样本复杂度,指出确定性瞬态和所需预热时间仍可能依赖于最小Bellman目标密度。

原文 · arXiv cs.LG

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes $α_t=c(t+1)^{-a}$ with $a\in(1/2,1)$, the leading last-iterate fluctuation is of order $\widetilde O\bigl(T^{-a/2}/\sqrt{1-γ}\bigr)$ and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order $m^{-1}$ in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.