研究人员发布了MR-JEPA模型,能处理心脏MRI多序列视频数据,在心脏功能评估和疾病检测上超越现有方法。
MR-JEPA是心脏磁共振成像(CMR)的自监督视频基础模型,通过tubelet标记化、时空掩码增强和2D CMR基础模型初始化扩展到3D时空输入。该模型在10,505名患者的多序列数据上进行预训练,涵盖cine、LGE和mapping序列。在六项下游任务评估中,MR-JEPA在五项回归任务上表现优于其他方法,包括LV EF MAE为4.79%(r=0.764)和GLS MAE为1.87(r=0.805),应变任务上比基线减少21-27%的MAE。
MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI
Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep learning approaches process individual 2D slices, discarding this context. We present MR-JEPA, a self-supervised video foundation model for CMR that extends LeJEPA to 3D spatiotemporal inputs through tubelet tokenization, spatiotemporal masking augmentation, and initialization from a 2D CMR foundation model. Unlike prior CMR video models limited to cine data, MR-JEPA is pretrained on multi-sequence data (cine, LGE, mapping) from 10,505 patients across two centers without annotations. We evaluate the frozen encoder on six downstream tasks using a unified multi-view gated attention architecture: LV ejection fraction, RV ejection fraction, three myocardial strains (GLS, GCS, GRS), and four-class disease detection. MR-JEPA outperforms other compared methods on all five regression tasks, including both a domain-specific CMR model pretrained on more data with text supervision and a natural-video foundation model, achieving an LV EF MAE of 4.79% (r =0.764) and a GLS MAE of 1.87 (r=0.805), with 21-27% MAE reductions over baselines on strain tasks. For disease detection, MR-JEPA achieved a macro AUG of 0.868, remaining competitive with the domain-specific baseline despite using a fully self-supervised pretraining objective. These results demonstrate the potential of a unified video encoder for robust, multi-view utilization of diverse CMR sequences in clinical cardiac quantification and diagnosis.