Wasserstein Policy Learning 用于分布结果型离线策略学习

Wasserstein Policy Learning for Distributional Outcomes

精选理由

这篇论文把因果推断中的离线策略学习扩展到了分布结果,用Wasserstein重心定义奖励,并给出了严格的统计保证,和传统均值策略学习不同,适合做理论研究的参考。

AI 摘要

该论文研究离线策略学习中结果变量为分布的情况,将每个潜在结果视为概率测度,并通过 Wasserstein 重心下的效用函数定义奖励。论文基于 IPW 和 Doubly Robust 估计量建立了统计保证,证明了有限样本后悔率的领先项为 O~(√(N-dim(Π)/N))。在一维 Wasserstein 设定下,后悔率仍由策略类复杂度主导。另外提供了极小化下界,证明了对 N 和 N-dim(Π) 的领先依赖的紧致性。

原文 · arXiv cs.LG

Wasserstein Policy Learning for Distributional Outcomes

Offline policy learning has received growing attention in causal inference. The primary objective is to learn a policy (individualized treatment rule) as a mapping from covariates to treatment that maximizes the empirical welfare defined as the mean of scalar-valued potential outcomes. In this paper, we study offline policy learning with distribution-valued outcomes, where each potential outcome is a probability measure on $\mathbb{R}$ and the reward is defined through a utility functional applied to the Wasserstein barycenter of induced outcome distributions. We establish statistical guarantees for the policy learning framework based on both Inverse Probability Weighting (IPW) and Doubly Robust (DR) estimators. By handling the challenging uniform deviation over the product of the combinatorial policy class and the infinite-dimensional quantile domain, we prove that the finite-sample regret has leading dependence $\widetilde{\mathcal{O}}(\sqrt{\mathrm{N\text{-}dim}(Π)/N})$. In the one-dimensional Wasserstein setting and under the stated regularity conditions, the leading regret rate is still governed by the policy-class complexity. Moreover, we provide a minimax lower bound establishing the sharpness of the leading dependence on $N$ and $\mathrm{N\text{-}dim}(Π)$.