做可解释性研究的团队终于有了一个统一的理论框架,能系统设计方法而非拼凑碎片,建议关注论文中的对称性和约束推导部分。
随着AI模型日益复杂,可解释性成为理解、调试和控制模型的关键工具,但该领域缺乏通用理论来演绎设计可解释方法,导致文献碎片化和评估标准不一致。为此,研究者提出了标准可解释模型(SIM),这是一种基于拉格朗日力学的通用理论,能从用户对可解释性的前提假设出发,系统推导出对称性和约束,进而构建拉格朗日函数,其最小值对应最优可解释模型。通过调整不透明模型参数或编译约束到可解释架构,可达到最小值。实验表明,SIM能识别并解决传统、概念和机制可解释性方法的局限性,揭示未充分探索的研究方向,并指导核心编程接口设计。该理论还为可解释性课程提供教学基础,有望改变该领域长期碎片化的现状。
The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics
As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, interpretability lacks general theories to deductively design interpretable methods. This gap between theories and methods results in a fragmented literature and inconsistent evaluation protocols. To fill this gap, we introduce the Standard Interpretable Model (SIM), a general theory grounded in Lagrangian mechanics that enables the deductive design of interpretable methods. Specifically, the SIM summarises, in a set of premises, what interpretability is for a target user. From these premises, the SIM systematically derives interpretability symmetries and corresponding constraints, which shape the landscape of a Lagrangian whose minima correspond to optimal interpretable models. To reach the minima, one can either update the parameter values of an opaque model to make it more interpretable or compile constraints into an interpretable architecture. We empirically show that the SIM identifies and solves limitations of existing methods (including traditional, concept-based, and mechanistic interpretability), highlights underexplored research directions, and informs the design of core programming interfaces. Beyond being a research method, the deductive nature of the SIM offers pedagogical grounding for interpretability curricula and may shift the scientific community's perspective of a discipline that has long been fragmented.