论文

EchoJEPA 超声预训练迁移到肺部超声结核筛查的研究

Wrong Organ, Right Physics: Transferring Echocardiography Pretraining to Lung Ultrasound for Tuberculosis Screening

精选理由

研究团队把心脏超声预训练模型搬到肺超声筛查结核,结果预训练领域影响几乎为零,一个简单特征标准化技巧反而涨了 1.23 个百分点,做医疗影像的可以看看。

该研究对比了 17 个编码器,考察超声心动图预训练能否迁移到标注数据稀缺的肺部超声(LUS)结核筛查任务。预设对比中,EchoJEPA-L 与 V-JEPA2-L 仅差 -0.16 个百分点(p=0.926),各编码器整体差距仅 2.50 个百分点,低于 2.71 的测量分辨率。真正起作用的是特征标准化,使 17 个编码器平均提升 +1.23 个百分点(p=1.5×10⁻⁵)。测试集上最佳编码器较基线提升最多 +2.57 个百分点 AUROC,90% 敏感度下特异度从 60.3% 提高到 79.3%。

原文 · arXiv cs.LG

Wrong Organ, Right Physics: Transferring Echocardiography Pretraining to Lung Ultrasound for Tuberculosis Screening

Lung ultrasound (LUS) is attractive for tuberculosis (TB) screening at primary-care level, but labelled cohorts are small. Echocardiography carries no such constraint, while sharing the same underlying ultrasound imaging physics, signal processing and B-mode appearance as LUS. We ask whether an encoder pretrained on that high-resource ultrasound domain carries representations that remain usable in the low-resource one. Only the encoder varies, across seventeen encoders spanning three architecture families. Among them, a latent-predictive video encoder pretrained on generic video (V-JEPA2-L) and its echocardiography counterpart (EchoJEPA-L) differ in pretraining corpus alone. The choice among these encoders does not resolve the classification, the whole family spanning 2.50 percentage points against a measurement resolution of 2.71. What moves the task instead is feature conditioning. Standardising the features between the encoder and the classifier improves all seventeen encoders by a mean of +1.23 percentage points at $p=1.5\times10^{-5}$. On the held-out test set every encoder selected on the development folds stands above the baseline system by up to +2.57 percentage points of area under the receiver operating characteristic curve (AUROC), and specificity at 90% sensitivity reaches 79.3% against 60.3%. The contrast specified in advance, EchoJEPA-L against V-JEPA2-L, measures -0.16 percentage points at $p=0.926$. We therefore find no evidence that shared ultrasonic physics alone makes echocardiography a more productive pretraining corpus than generic video, and any advantage, if present, is smaller than this cohort can resolve. The video encoders receive replicated still images, however, so whether this absence of an effect reflects the pretraining domain or a video encoder applied to static frames cannot be separated. The limiting factor is the labelled cohort rather than the encoder.