SMILE:医疗诊断自解释多模态信息瓶颈
SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis
医学AI可解释性新方法SMILE,在多模态诊断中表现优异,比现有方法准确率高9.1%,还能解释诊断依据。
SMILE是一种基于信息框架的自解释多模态诊断方法,在iCTCF数据集上实现9.1%的准确率提升。该方法通过联合优化预测性能和模态特定可解释性,识别各模态中贡献诊断决策的最具信息量元素。实验表明,该方法在异构模态医疗数据集上保持强诊断性能,并提供透明且模态感知的特征相关性见解。
SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis
Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in safety-critical clinical decision-making. Most existing explainability methods in healthcare operate in a post-hoc manner and are predominantly designed for unimodal data, which limits their applicability in increasingly prevalent multimodal diagnostic settings. This paper addresses the problem of self-explainable multimodal diagnosis by formulating it within the information bottleneck (IB) framework. We propose a unified learning paradigm that jointly optimizes predictive performance and modality-specific explainability by identifying the most informative elements inside each modality that contribute to diagnostic decisions. To enable tractable and stable optimization, we employ a matrix-based Renyi's $α$-order entropy functional under the assumption of sufficiently expressive encoders. Extensive experiments on representative medical datasets spanning heterogeneous modalities demonstrate that the proposed method consistently achieves strong diagnostic performance, including an absolute accuracy improvement of 9.1 percentage points on the iCTCF dataset. Moreover, the learned explanations provide transparent and modality-aware insights into feature relevance, thereby improving both the explainability and generalization.