压力下的逻辑判断:用学习软前缀诊断三段论稳定性

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

精选理由

这篇论文用软前缀测试了三段论推理,发现模型答案很容易被带偏,Qwen和Gemma的稳定性差异很大,值得做推理研究的读一下。

AI 摘要

该论文通过在精确标注的三段论推理基准上添加软前缀,测试了Qwen3.6-35B-A3B MoE、Qwen3-8B和Gemma 4 31B三个模型在逻辑判断中的稳定性。软前缀能覆盖大多数正确答案,并在16组模型-方向-分割比较中,比随机控制前缀的翻转率高出37到99个百分点。Qwen3.6 MoE在不同措辞和提示变化下的翻转率维持在72%至90%,而Gemma的有效前缀翻转率为54%至56%,远低于匹配随机前缀的不到1%。诊断测试表明,主导效应是对某一种答案含义的广泛偏好,而非固定的符号强制或跨任务可靠转移的运算,且这种偏差形式因模型而异。

原文 · arXiv cs.AI

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical form and interface. By studying which prefixes succeed and how their effects generalize, we characterize how learned contextual pressure can override correct judgments and expose limits in a model's logical stability. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, learned prefixes redirect many correct answers and remain effective across unseen forms and interface changes. In repeated tests with Qwen3.6 MoE and Gemma, they outperform paired random controls in all 16 model--direction--split comparisons by 37 to 99 percentage points. Qwen3.6 MoE flip rates remain between 72% and 90% across wording and prompt changes, while Gemma validity prefixes retain 54% to 56% flip compared with less than 1% for matched random prefixes. Diagnostic tests show that the dominant effect is a broad preference for one answer meaning rather than fixed-symbol forcing or a logical operation that transfers reliably between tasks. The form of this bias differs across models. In both Qwen models, simple score models often predict which judgments will flip but not how far their margins will move, whereas Gemma's overall response is more closely approximated by the same models. These results show that the dominant behavioral effect of successful soft prefixes is a broad answer preference, while the remaining response reveals substantial model-specific differences in logical stability.