想不靠人工标注就给模型打分?这篇论文用数学定理把理性检验变成可计算惩罚,还给了三个现成公式。
这篇论文把决策论中的表示定理用于AI评估:检查模型行为是否满足公理,不需要外部标签或人类反馈。作者用德菲内蒂定理做概率一致性检验、用阿夫里亚特定理检验偏好理性、用Echenique和Saito(2015)定理检验主观期望效用。每个检验都给出连续惩罚项,当行为可被理性化时惩罚为零。由于公理是充要条件,通过检验的模型在理性标准下无法被同一数据上的其他测试拒绝。这些惩罚不限制具体目标函数,因此可与其他评估和训练信号互补。
Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems
Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well-defined objective. I argue that this ``if and only if'' structure provides a potentially useful foundation for label-free evaluation and regularization of LLMs and other AI systems. Axiom compliance can be checked from the model's own responses to synthetic choice problems, with no external labels or human feedback, and the penalties are readily computable. Because the axioms are necessary and sufficient, the resulting checks exhaust the implications of the relevant rationality standard for the elicited data: a model that passes cannot be rejected on rationality grounds by any further test of the same data. I discuss three instantiations: probabilistic coherence via a theorem of de Finetti, preference rationality via Afriat's theorem, and subjective expected utility via a theorem of Echenique and Saito (2015), each yielding a continuous penalty that is zero whenever behavior can be rationalized. Since coherence does not restrict which objective rationalizes behavior, these penalties complement rather than replace other evaluation and training signals.