新框架显式校正 LLM 评审偏差,用更少对比恢复可靠排名
Accounting for Bias Enables Sustainable LLM Evaluation
做模型评测的朋友看看这篇:它把 LLM 当裁判的位置偏好、啰嗦偏好这些偏差直接建模掉,能省大量对比次数。
一篇 arXiv 论文指出,LLM-as-a-judge 评测中 position bias、verbosity bias、judge severity 和 self-enhancement 四类偏差无法靠增加对比次数消除,现有排行榜靠堆数据的做法在统计上不成立。论文提出统一潜变量框架,同时建模成对与序数数据并显式校正这些混杂因素,用明显更少的对比次数即可恢复可靠排名。该框架的拟合计算成本相对一轮 LLM 推理可忽略不计。
Accounting for Bias Enables Sustainable LLM Evaluation
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.