论文精选

Meta新论文:LLM裁判被质疑时判决翻转率25%-91%

精选理由

Meta用Wiggle框架折腾了9个前沿模型,发现LLM裁判一被追问就翻案,最高91%翻转,这评测挺扎心。

Meta新论文提出Wiggle Framework,压力测试LLM裁判的判决稳定性。他们用9个前沿模型在14项评判任务上测试了重新提示、单次质疑和持续施压三条轴线。结果显示所有模型都会改变判决,静态反驳下翻转率为25%-71%,面对对抗性说服者则达62%-91%。这种压力导致的判决改变几乎总是偏离地面真值,基线陪审团多数强度是预测哪些案例会翻转的最佳指标。

原文 · elvis

Brilliant new paper from Meta. LLM judges get validated on accuracy against golden data. That says nothing about whether the verdict survives when questioned. The Wiggle Framework stress-tests 9 frontier models across 14 judging tasks along three axes, stability under re-prompting, stability under a single challenge, and stability under sustained pressure. They find that every model wiggles. Verdicts flip 25 to 71% of the time under static pushback, and 62 to 91% against an adversarial persuader. Pressure that changes a judge's verdict is almost always net-corrupting against ground truth. Baseline jury majority strength turns out to be the best single-shot predictor of which items will move. Paper: arxiv.org/abs/2608.12645 Track more trending AI papers in our academy: academy.dair.ai 💬 0 🔄 1 ❤️ 2 👀 482 📊 1 ⚡