论文精选

一种新算法提供无分布假设的安全保障

Conformal Policy Learning with Distribution-Free Safety Guarantees

精选理由

朋友,有个新算法叫conformal policy learning,它能在做决策时保证不伤害人,特别适合高风险领域,比如医疗或公共政策。

这篇论文提出了一种名为‘conformal policy learning’(CPL)的新方法,它通过使用可观测的代理变量和选择性校准来控制个体受到伤害的概率。在随机实验中,该方法在标准交换性条件下,无需任何结果建模假设,就能在指定水平上提供有限样本的安全保障。此外,当结果模型被一致估计时,CPL在满足安全约束的情况下实现了渐近最优福利。

原文 · arXiv cs.LG

Conformal Policy Learning with Distribution-Free Safety Guarantees

Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose \textit{conformal policy learning} (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.