放射科医生和医学影像 AI 研究者终于有了一个能真正做前后对比和参考病例检索的框架——MedReCo 在 12 项检索任务中全胜,做临床 AI 落地的团队值得关注。
该研究提出了一个实体感知的跨图像比较推理框架 MedReCo,用于解决放射科实践中依赖前后对比和参考病例的诊断需求。研究构建了 MedReCo-DB 大规模数据集,包含来自 8 家机构、4 个国家、7 种影像模态的 69 万张图像,并将报告分解为解剖结构、异常发现和病理条件。基于此,开发了用于可控检索的 MedReCo 编码器和用于生成式比较解读的 MedReCo-VLM 视觉语言模型。在内部、外部和跨中心评估中,MedReCo 在 12 项内部检索设置中均取得最高 Recall@1,外部检索平均提升 6 个百分点;MedReCo-VLM 在比较生成评估中全面最优,纵向随访准确率提升 14.5-46.5 个百分点(胸片)和 13.0-27.9 个百分点(CT)。这表明实体感知的比较推理可从常规临床数据中大规模学习,为医学影像 AI 提供更贴近临床的范式。
A Vision-language Framework for Comparative Reasoning in Radiology
Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological practice, where diagnosis and follow-up rely on comparison across prior studies and analogous reference cases. Here we formulate radiological comparison as an entity-aware cross-image reasoning problem and introduce a framework that supports both reference-case retrieval and temporal comparative interpretation. We construct MedReCo-DB, a large-scale comparative imaging resource derived from routine image-report pairs, comprising more than 690,000 images from over 160,000 patients across eight institutions, four countries and seven imaging modalities. Reports are decomposed into anatomical structures, abnormal findings and pathological conditions to provide supervision for entity-conditioned retrieval and comparative visual question answering. Using this resource, we develop MedReCo, an entity-aware visual encoder for controllable retrieval of clinically analogous cases, and MedReCo-VLM, a vision--language extension for generative interpretation of interval change. Across internal, external and cross-center evaluations, MedReCo achieved the highest Recall@1 in all 12 internal retrieval settings and improved external retrieval by a mean of 6.0 percentage points. In clinically confusable differential groups, it consistently outperformed the strongest baselines. MedReCo-VLM achieved the best performance across all comparative generation evaluations and improved longitudinal follow-up accuracy by 14.5-46.5 percentage points on chest radiographs and 13.0-27.9 percentage points on CT. These findings suggest that entity-aware comparative reasoning can be learned from routine clinical data at scale and may provide a more clinically aligned foundation for medical imaging AI.