Auditing Question-Order Effects in LLMs with the QQ Equality

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

精选理由

这篇论文用QQ等式定量审计LLM的顺序效应,还给出了理论机制和实用流水线,适合对模型行为一致性感兴趣的人读。

AI 摘要

该论文将量子问题顺序模型中的QQ等式发展为审计自回归大语言模型(LLMs)顺序判断的标准。理论表征了满足QQ等式的机制类别,包括边际独立核和极性依赖重复家族。方法上开发了预指定审计流水线,结合最坏情况鲁棒性包络、采样一致性抽查和饱和诊断。在开放权重指令微调模型上进行的初步实验中,所有预指定的健康门通过,但17/18和7/8项目对饱和(近确定性),没有项目通过残差上下文认证。

原文 · arXiv cs.AI

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

Human survey respondents exhibit question-order effects that satisfy the QQ (quantum question) equality, an a priori, parameter-free prediction of the projective quantum question-order model. We develop the QQ equality into an audit criterion for sequential judgments of autoregressive large language models (LLMs). Theoretically, we characterize which mechanism classes satisfy it robustly: marginal-independent kernels satisfy QQ iff all four mismatch transition rates coincide (a class containing the 2D rank-1 projective model with a fixed measurement pair under state variation); a polarity- and position-dependent repetition family is characterized by an exact cross-symmetry condition with closed-form violations; QQ-satisfying behaviors are closed under order-matched mixing; and the rank-2 Contextuality-by-Default criterion translates into audit coordinates as $|\qQQ|\le\OSS$, where $\OSS$ (the order-sensitivity score) totals the order sensitivity of the two marginals. Methodologically, we develop a pre-specified, audit-logged pipeline applicable to any model exposing next-token log-probabilities; it combines worst-case robustness envelopes, sampling-consistency spot checks, full label counterbalancing, and a saturation diagnostic. Empirically, in a first-signal pilot on an open-weight instruction-tuned model under two framings, all pre-specified health gates passed, yet 17/18 and 7/8 item pairs, respectively, were saturated (near-deterministic), and no item was certified residually contextual. Forced-binary next-token log-probabilities were thus inadequate for distribution-level QQ audits under the tested model and prompting conditions; we recommend pre-specified saturation diagnostics whenever next-token distributions are treated as survey-response distributions.