想系统搞懂CAM方法演进,这篇57篇论文的综述把梯度法、Transformer归因、CLIP/DINO/SAM时代的新做法都梳理清楚了,还指出了各方法留下的坑。
该综述聚焦类激活映射(CAM)这一可解释视觉技术家族,梳理了自2016年首个CAM提出以来的57篇方法论文。文中按归因机制、架构依赖和评估目标构建分类体系,涵盖梯度法、无梯度评分、高分辨率上采样、弱监督定位、Transformer token归因以及基于CLIP、DINO、SAM的方法。综述指出研究趋势正从单层低分辨率类别分数解释转向多层、概率化、token感知和基础模型感知的比较式解释。评估协议仍碎片化,忠实度、定位、鲁棒性、计算成本和人类信任缺乏统一标准。
Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.