IBM这篇研究说明,看榜单选模型可能被题目措辞骗了,强模型反而更怕改写。
IBM发布BenchDrift研究,生成基准问题的语义等价变体来测试模型鲁棒性。结果显示,弱模型从改写中获益多于损失,而强模型损失远大于获益。在GSM8K、MMLU和MATH-Hard上测试的8个模型中,措辞敏感性随模型增强不降反增。该研究表明,顶尖模型的基准分数高度依赖题目措辞,而非纯粹能力。论文已发布于arxiv.org/abs/2608.11694。
Interesting new research from IBM. If you pick models from benchmark deltas, some of that delta bel...
Interesting new research from IBM. If you pick models from benchmark deltas, some of that delta belongs to the phrasing rather than the model. BenchDrift generates meaning-preserving variations of benchmark problems along linguistic, referential, pragmatic, and structural axes, holding the answer fixed, then measures how often correctness flips. Phrasing sensitivity does not fade as models improve. It changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain, so the top models on a benchmark are the ones whose scores depend most on the wording they happened to receive. Fragility also belongs to the rephrasing. Across eight models on GSM8K, MMLU, and MATH-Hard, they largely agree on which rephrasings cost the most correct answers even while differing in how much they drift overall. Rephrasing breaks answers models were confident about, whether the problem gets shorter or longer. Paper: arxiv.org/abs/2608.11694 Track more trending AI papers in our academy: academy.dair.ai 💬 6 🔄 8 ❤️ 35 👀 5964 📊 15 ⚡