cs.CVSep 30, 2026

Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models

Authors: Bangwei Guo, Xiao Chen, Boris Mailhe, Jia Yao, Yiqing Wang, Ankush Mukherjee, Yikang Liu, Zheyuan Zhang, +3 more

Organizations: Rutgers University, NJ, USA · United Imaging Intelligence, Boston, MA, USA · University of Texas Southwestern Medical Center, TX, USA · Duke University, NC, USA

Abstract

Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians interpret these images by identifying cardiac structures and focusing on the regions relevant to each clinical question, motivating anatomically guided vision-language models (VLMs). Yet CMR-specific supervision for anatomical localisation and clinical question answering remains limited. To address this gap, we investigate fine-grained CMR visual question answering through anatomical grounding and guided attention. We construct 128,915 anatomical-grounding and 42,799 clinical QA pairs across short-axis cine, late gadolinium enhancement, and long-axis cine. These datasets support anatomical recognition, localisation, and clinical assessment without requiring paired reports for individual training images. To help the model learn where to look, we introduce Cardiac Anatomy-Routed Attention (CARA), which selects predicted anatomical priors according to the question and guides decoder attention with learned task-specific strengths. Combining anatomical grounding pretraining with CARA yields our model, CARA-VL. Experiments demonstrate CARA-VL's strengths in clinical assessment and regional localisation across CMR imaging settings, with promising generalization to an external clinical cohort. Together, our data and method provide a practical framework for studying and advancing cardiac visual understanding in VLMs. We will release the QA data derived from public datasets upon publication.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

    Oct 1, 2026Kunyang Li, Hai Nguyen, Joshua Lowe +6Cardiac Magnetic Resonance ImagingFilms

  2. CMRVision: A Foundation Model for Cardiac MR Image Analysis

    Sep 1, 2026Athira J. Jacob, Puneet Sharma, Daniel RueckertCardiac Magnetic Resonance ImagingEchocardiography

  3. CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

    May 28, 2026Zixian Su, Hongkai Zhang, Fan Gao +12Cardiac Magnetic Resonance ImagingMedical Vision-Language Models