论文

Conformal Prediction 集合大小可作为信息增益度量的理论证明

Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective

精选理由

用 conformal prediction 做不确定性估计的朋友看看,这篇论文证明了预测集变小就等于获得了信息,还给了 11 个实验验证。

arXiv 论文为 conformal prediction 的预测集大小提供了信息论基础,引入了一族基于集合大小和覆盖率的广义信息度量,并证明 Shannon 互信息可用这些度量精确的积分表示。在标准分类场景中,额外信息带来的集合缩小幅度被该族度量上下界夹住,且满足数据处理不等式(误差限于有限样本校准和模型误差项)。作者在 11 个分类设置上验证了理论,并在贪心特征选择实验中发现集合缩小量与 Shannon 互信息可能给出不同的特征排序。

原文 · arXiv cs.LG

Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective

Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty, the information-theoretic basis for this interpretation remains poorly understood. In this work, we provide such a foundation using a decision-theoretic generalization of entropy tailored to set-valued prediction. In particular, we introduce a family of generalized information measures based on the size and coverage of conformal prediction sets. Notably, Shannon mutual information admits an exact integral representation in terms of these measures. We then show that, in standard classification settings, the reduction in conformal set size from additional information (i) is sandwiched between calibration-dependent members of this family and (ii) obeys a data processing inequality, both up to finite-sample calibration and model error terms. Together, our results formally relate conformal prediction to classical information-theoretic quantities and justify using set-size reduction as an information gain metric. Empirically, we validate our theory across 11 classification settings and show that set-size reduction and Shannon mutual information can rank features differently in a greedy feature selection experiment.