这篇论文用数学告诉你,LLM自报置信度不靠谱时,按置信度排序审计可能比随机还差,而且越省预算越危险,值得做AI治理的人看看。
一篇arXiv论文研究单人对N个LLM智能体舰队在每轮预算B≪N次审计下的监督问题,依据自我报告置信度排序,但置信度可能被对抗性错误校准且误差相关。作者用双层高斯Copula建模,找到了错误校准阈值δ*,超过该阈值后按置信度排序审计比随机审计更差。两个先验预期被推翻:预算越少δ*反而越高,跨家族相关性并不低——共享难度主导了谱系。五个开源模型置信度近乎恒定、无操作价值,点估计处于或超过翻转点但置信区间跨越;一个专有模型有信息量且落在阈值以下。论文还给出了“空洞监督”的定量判据,并通过历史轨迹回放验证了排序。
One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold $δ^*$ past which confidence-ranked auditing is \emph{worse} than random. Two a-priori expectations reverse: $δ^*$ \emph{rises} as the budget shrinks, and cross-family correlation is not low---shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for \emph{vacuous} oversight, and replaying policies on recorded traces confirms the ordering.