基于视觉语言模型的心脏消融规划LGE-MR图像质量评估

Toward Vision Language Model-based Assessment of Clinical Quality and Usability of LGE-MR Images for Cardiac Ablation Planning

精选理由

这篇论文提出了一种基于视觉语言模型的心脏消融规划LGE-MR图像质量评估方法,通过两个阶段的模型预测图像质量,有助于提高消融过程的安全性。与现有方法相比,InternVL2和DeepSeek在准确性和临床可用性方面表现出色。

AI 摘要

LGE心脏MRI在心房颤动患者中用于左心房纤维化评估和消融规划,图像质量对消融过程至关重要。本研究提出一种两阶段视觉语言模型(VLM)框架,用于左心房LGE-MRI的临床图像质量评估。第一阶段,微调的VLM生成结构化的放射学风格质量报告,预测五个放射科医生定义的标准:噪声、运动伪影、左心房边界准确性、肺静脉区域准确性和过度分割严重程度。第二阶段,基于GPT的推理模块将预测的质量和报告映射到结构化的质量评分和二进制临床可用性决策。使用60个标注的图像切片-文本对进行基准测试,InternVL2在标准级别上达到最高准确率(平均ACC=0.65,PLCC=0.79),DeepSeek在临床可用性上达到完美的一致性(准确率=1.00,kappa=1.00)。

原文 · arXiv: DeepSeek

Toward Vision Language Model-based Assessment of Clinical Quality and Usability of LGE-MR Images for Cardiac Ablation Planning

LGE cardiac MRI is widely used for left atrial fibrosis assessment and ablation planning in atrial fibrillation patients as knowledge of fibrotic tissue regions identified from LGE-MRI is critical for catheter ablation. Often, poor quality images used during ablation planning can cause mis-localization of ablation targets, directly impacting procedure safety and outcome. The decision of whether a scan meets the minimum quality threshold for ablation planning is currently made informally by the reviewing radiologist and is not captured by any automated system, yet it is arguably the most safety-critical output of the image quality assessment (IQA) process. However, variations in image quality caused by noise, motion artifacts, and poor boundary definition significantly compromise the reliability of downstream segmentation and clinical decision-making tasks. Manual quality assessment by expert radiologists is subjective and difficult to scale, while existing automated methods produce scalar scores without interpretable clinical reasoning. In this work, we propose a two-stage vision language model (VLM) framework for clinically grounded image quality assessment of left atrial LGE-MRI. In the first stage, a fine-tuned VLM generates structured radiology-style quality reports predicting five radiologist-defined criteria: Noise, Motion Artifact, LA Boundary Accuracy, PV Region Accuracy, and Under-segmentation Severity. In the second stage, a GPT-based reasoning module maps the predicted quality and reports to a structured quality scores and binary clinical usability decision for ablation planning. We curate a dataset of 60 annotated image slice-text pairs from 20 patients and benchmark four state-of-the-art VLM architectures. InternVL2 achieves the highest criterion-level accuracy (Avg ACC=0.65, PLCC=0.79), while DeepSeek achieves perfect clinical usability agreement (Acc=1.00, kappa=1.00).