多组均值估计主动学习的复杂度度量:方差局部曲率

A Complexity Measure for Active Learning in Multi-group Mean Estimation

精选理由

新复杂度指标VLC揭示主动学习难度来源

AI 摘要

本文研究多组均值估计主动学习的 max-risk 目标:在 d 个臂中分配 T 次采样以最小化最坏情况不确定性指数 max σ_k²/n_k。作者提出局部最小最大化框架,证明首个针对该目标的一般下界,将难度分解为预算项、异质性指数和模型相关复杂度度量 VLC。VLC 可重参为方差-费希尔信息,并为常见分布族给出闭式解。与现有上界对比,在广泛场景下接近最优(对数因子内),但高异质性实例存在系统差距。

原文 · arXiv cs.LG

A Complexity Measure for Active Learning in Multi-group Mean Estimation

We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}σ_k^2/n_k$, where $σ_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled. We develop a local minimax framework and prove the first general lower bound for this objective, valid for any finite-variance hypothesis class. The bound separates difficulty into three orthogonal factors: a \emph{budget} term, a \emph{heteroscedasticity} index measuring how unevenly the uncertainty is spread across arms, and a model-dependent complexity measure, the \emph{Variance Local Curvature} ($\mathrm{VLC}$), which captures how much information a local change of variance creates inside the hypothesis class. For smooth classes, the $\mathrm{VLC}$ is a reparametrization of a variance--Fisher information, with closed-form values for common families. Benchmarking against the strongest available upper bound shows near-optimality up to logarithmic factors in broad regimes, and pinpoints a systematic gap in highly heterogeneous instances. Our proof introduces two key ingredients: a loss-induced $\ell_1$ geometry on the decision space, and a representation-based instance generator that reduces hard-instance construction to an explicit random matrix calculation.