正则化为何有效:对抗模仿学习的快速率理论证明
Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
GAIL 这些老方法里加正则化为什么效果好,这篇终于给出了数学证明,还配套了新算法,做强化学习理论的朋友可以看看。
论文研究对抗模仿学习(AIL),指出 GAIL 和 LS-IQ 等方法中常用的奖励正则化与熵策略正则化此前缺乏有限样本理论支撑。作者提出 Dually Regularized AIL 算法,将 KL 策略正则化与按专家和学习者占用度加权的二次奖励惩罚结合。在 K 次在线交互与 N 条专家轨迹下,正则化模仿差距被证明为 O(1/K + 1/N)。据作者所称,这是首个在随机专家情形下同时在专家演示和在线交互两个维度达到 O(1/ε) 样本复杂度的算法。
Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a $\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$ bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}ε\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.