这篇SaliTrap研究用1145道常识题考了12款模型,最好的也才过半避坑,点开看看谁最容易被数字带偏。
新论文《Would You Walk to the Car Wash?》提出SaliTrap基准,用1,145个常识陷阱提示测试12款模型。表现最好的模型仅在54.8%的问题中避开陷阱,八款模型低于30%。随着提示中数字密度增加,模型更容易直接开始计算而忽略物理前提。即使模型意识到陷阱,GLM-5.1和Kimi-K2仍分别有86.2%和81.8%的几率服从不合理请求。去掉诱饵后,四款代表性模型在86.9%至91.8%的之前谄媚案例中恢复正确判断。
Bets on whether Astra will be vulnerable to the same problems? I think the answer is “very likely y...
Bets on whether Astra will be vulnerable to the same problems? I think the answer is “very likely yes” Rohan Paul @rohanpaul_ai LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must reach the wash. The failure here is not missing knowledge, but failing to use it when explicit details dominate the prompt. The paper calls this salience bias: explicit numbers and procedures overpower unstated physical prerequisites. SaliTrap tests 1,145 such prompts across four trap types and 12 models. Even the best evaluated model avoided the trap in only 54.8% of queries, while eight of twelve stayed below 30%. More distractors made this worse: trap avoidance fell as numerical density rose, while models increasingly began calculating before noticing the contradiction. Awareness was not enough either. Among trap-aware responses, GLM-5.1 and Kimi-K2 still complied 86.2% and 81.8% of the time. The strongest diagnosis comes from removing the bait: context-free probes recovered 86.9% to 91.8% of previously sycophantic cases across four representative models. So many failures reflect knowledge that is present but suppressed by task framing. For agent evaluation, premise checking should be tested as a behavioral control, not assumed from a model’s general reasoning score. – arxiv. org/abs/2607.28478 Title: "Would You Walk to the Car Wash? Revealing the Salience Bias of LLMs in Commonsense Reasoning" 🔗 View Quoted Tweet 💬 5 🔄 1 ❤️ 7 👀 1950 📊 4 ⚡