基于机器学习的咳嗽模型为何无法泛化:结核病筛查的系统跨数据集评估

Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening

精选理由

这篇论文详细评估了咳嗽结核病模型的泛化能力,揭示了数据收集和设备差异对模型泛化能力的影响,为结核病筛查模型的开发提供了重要参考。

AI 摘要

咳嗽声学在非侵入性结核病筛查中具有潜力,但机器学习(ML)模型是否捕捉到疾病相关的声学特征或数据收集的伪影尚无定论。研究人员评估了经典ML和深度学习(DL)咳嗽结核病分类器在三个独立数据集上的跨数据集泛化能力。尽管在数据集内部表现良好(ROC-AUC高达0.755±0.056),但两种方法都无法泛化,外部性能通常低于0.6,表明数据可能存在局限性。此外,音频表示由录音设备和数据集组织,而不是结核病状态,预测结核病概率与CODA国家层面的流行率一致,设备不匹配会降低迁移能力,而设备多样化的训练则会提高它。此外,临床变量基线泛化更一致(ROC-AUC 0.655-0.711),表明获取特定变量的可变性是泛化能力差的主要驱动因素。数据集内部的高性能不足以保证泛化能力。在咳嗽结核病模型临床应用之前,外部验证是必要的。

原文 · arXiv cs.AI

Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening

Cough acoustics are promising for non-invasive tuberculosis (TB) screening, yet whether machine learning (ML) models capture disease-related acoustics or artifacts of data collection remains unresolved. We evaluated the cross-dataset generalizability of classical ML and deep learning (DL) cough-based TB classifiers across three independent datasets. Despite moderate within-dataset performance (ROC-AUC up to $0.755 \pm 0.056$), both pipelines fail to generalize, with external performance frequently below 0.6, indicating a possible limitation of the data. We further observed audio representations are organized by recording device and dataset rather than TB status, predicted TB probability tracks country-level prevalence in CODA, and device mismatch degrades transfer while device-diverse training improves it. Additionally, a clinical-variable baseline generalizes more consistently (ROC-AUC $0.655 - 0.711$), indicating acquisition-specific variability is a stronger driver of poor generalizability than population shift. High within-dataset performance is not enough. External validation is essential before cough-based TB models are clinically ready.