研究模型选择需多少标签?通过证书和预算量化选择性预测
How Many Labels Does Model Choice Need? Certificates and Budgets for Selective Prediction
想了解模型选择时到底需要多少标签?这篇论文用数学方法量化了,比如K个候选模型时,平均需要56-57%的标签,比之前估计的20%要高很多。
论文提出用广义风险覆盖曲线下面积(AUGRC)量化模型选择所需标签数量。对于K个候选模型,当所有标签已知时,覆盖线性规划可确定最小标签数(证书大小),误差独立时该数量接近池的四分之一。在108个特征面板比较中,不同标签解决所有准确率选择,但无AUGRC选择。20%预算在96种情况下被排除,平均证书需求为56-57%。在10个预训练图像分类条件下,置信度选择读取68-91%的标签用于精确选择,在AUGRC容差为5×10⁻⁴时读取50-67%。
How Many Labels Does Model Choice Need? Certificates and Budgets for Selective Prediction
Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently. We quantify this requirement for the area under the generalized risk-coverage curve (AUGRC). A prelabel lower bound rules out insufficient budgets. With all labels known, a covering linear program bounds the minimum number of labels sufficient to fix the winner (the certificate size) within $K-1$ labels for $K$ candidates. For fixed $K$, independent uniform orders and identical predictions, the prelabel bound approaches one quarter of the pool. With iid Bernoulli errors independent of the orders, every exact acquisition policy reads almost all labels asymptotically, although a two-candidate certificate needs only half. Across 108 feature-panel comparisons on nine datasets, disagreement labels settle every accuracy choice but no AUGRC choice. A 20% budget is ruled out in 96 conditions; certificates need 56-57% on average. On ten conditions with pretrained image classifiers, confidence-score choice reads 68-91% of 10,000 labels for exact selection and 50-67% with AUGRC tolerance $5\times10^{-4}$. An exact stopping test works with any acquisition order. Together, these results link confidence ranks to label budgets and certified model comparison.