Meta用Wiggle框架折腾了9个前沿模型,发现LLM裁判一被追问就翻案,最高91%翻转,这评测挺扎心。
Meta新论文提出Wiggle Framework,压力测试LLM裁判的判决稳定性。他们用9个前沿模型在14项评判任务上测试了重新提示、单次质疑和持续施压三条轴线。结果显示所有模型都会改变判决,静态反驳下翻转率为25%-71%,面对对抗性说服者则达62%-91%。这种压力导致的判决改变几乎总是偏离地面真值,基线陪审团多数强度是预测哪些案例会翻转的最佳指标。
Brilliant new paper from Meta. LLM judges get validated on accuracy against golden data. That says ...
Brilliant new paper from Meta. LLM judges get validated on accuracy against golden data. That says nothing about whether the verdict survives when questioned. The Wiggle Framework stress-tests 9 frontier models across 14 judging tasks along three axes, stability under re-prompting, stability under a single challenge, and stability under sustained pressure. They find that every model wiggles. Verdicts flip 25 to 71% of the time under static pushback, and 62 to 91% against an adversarial persuader. Pressure that changes a judge's verdict is almost always net-corrupting against ground truth. Baseline jury majority strength turns out to be the best single-shot predictor of which items will move. Paper: arxiv.org/abs/2608.12645 Track more trending AI papers in our academy: academy.dair.ai 💬 0 🔄 1 ❤️ 2 👀 482 📊 1 ⚡