BiasFlow 发布:监测并缓解模型对虚假特征的依赖
BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance
研究模型偏置的可以看看:arXiv 上的 BiasFlow 同时给了监测质心几何的工具和 BFR 正则项,CelebA 压力测试里 WGA 从 40.7% 拉到 64.1%。
BiasFlow 是一套基于 hook 的诊断工具,可监测类属性质心对齐(IBMI)、类内质心分离(W-IBMI)和特征投影敏感度。配套的 BiasFlow Regularization(BFR)是一种有监督的类条件质心对齐惩罚项。在 UrbanCars 小规模基准上,加入 BFR 后最差组准确率(WGA)最多提升 26.0 个百分点。在 CelebA-Std 冻结骨干上加偏置数据训练新头的压力测试中,BFR+GroupDRO 将 WGA 从 40.7% 提升到 64.1%,同时 Male 探针准确率从 92.5% 降到 72.2%。合成水印 ImageNet 实验在匹配训练条件下将水印偏移准确率提升 23.0 个百分点。
BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance
Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class centroid separation (W-IBMI), and feature-projection sensitivity. IBMI is confounded by class-attribute correlation and is not a measure of causal feature reliance. We pair these diagnostics with BiasFlow Regularization (BFR), a supervised, composable class-conditional centroid-alignment penalty. W-IBMI verifies the quantity BFR optimizes; it is scale dependent and does not independently establish attribute removal. Across the reported small-scale benchmarks, adding BFR improves or preserves mean WGA, with gains up to +26.0 pp on UrbanCars. The principal independent stress test freezes CelebA-Std backbones and trains fresh heads on biased data: BFR+GroupDRO improves WGA from 40.7% to 64.1%, while Male probe accuracy decreases from 92.5% to 72.2%. Attribute information remains recoverable, and cross-task results are mixed. A controlled synthetic-watermark ImageNet experiment additionally improves watermark-shift accuracy by +23.0 pp under matched training. These results support evaluating centroid geometry and resistance to biased head retraining alongside WGA, within the tested protocols.