From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning
Authors: Jialu Pi, Yanan Ma, Weijie Chen, Owen Crystal, Shubham Trivedi, Stephen Xie, Anna Silverman, Matthew Stib, +3 more
Organizations: Dept. of Radiology, Mayo Clinic, Phoenix, AZ, USA · School of Computing and Augmented Intelligence, Arizona State University, Tempe, USA · Dept. of Radiology, Mayo Clinic, Phoenix, USA · Mayo Clinic, Phoenix, USA · Dept. of Cardiology, Mayo Clinic, Phoenix, USA
Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical data offers a scalable approach for identifying high-risk individuals before acute events occur. Although chest X-rays (CXRs) capture latent cardiovascular biomarkers and clinical histories provide complementary patient context, existing medical vision-language models are primarily optimized for radiology interpretation rather than prognostic reasoning. We propose a causal reinforcement learning framework for multimodal clinical reasoning that integrates CXRs and physician-authored clinical histories for opportunistic MACE prediction. The framework introduces (1) a role-decoupled dual-LLM architecture that separates reasoning from risk prediction, (2) a dual-action causal reinforcement learning policy for evidence selection and reasoning optimization, and (3) causal token pruning to learn compact multimodal representations. Evaluated on an internal cohort, an emergency department cohort, and the external MIMIC dataset, the proposed framework consistently outperformed unimodal baselines and state-of-the-art medical vision-language models, achieving AUROCs of 0.720, 0.760, and 0.845, respectively. It also substantially improved reasoning quality, achieving higher GREEN scores and higher expert preference while maintaining robust predictive performance across diverse patient populations.
Figures & tables
Fig. 1 : Proposed causal reinforcement learning framework with a multimodal LLM. (a) End-to-end training pipeline integrating multimodal inputs, multimodal token generation, causal reasoning, and reinforcement learning. Red arrows indicate gradient backpropagation and parameter updates. (b) Causal policy optimization, where reinforcement learning refines the reasoning policy using causal feedback to improve reasoning consistency and align generated reasoning with targeted clinical prediction objectives. LLM denotes large language model, and LM denotes language model.
Characteristic
Total
Mayo Internal
MIMIC External
Mayo ED External
P-value
Number of patients
22,573
15,802
4,616
2,155
p<0.001
Demographics
Age (years)
63.14 ± 17.34
64.16 ± 18.01
61.84 ± 17.33
58.38 ± 9.65
–
Female sex, n (%)
9,611 (42.6%)
6,454 (40.6%)
2,343 (50.8%)
832 (38.6%)
p<0.001
Male sex, n (%)
13,001 (57.6%)
9,431 (59.3%)
2,272 (49.2%)
1,323 (61.4%)
p<0.001
Race, n (%)
TABLE I : Baseline demographic and clinical characteristics of patients across the internal and external cohorts. Continuous variables are presented as mean ± standard deviation (or median [IQR], as appropriate), and categorical variables are presented as number (%). P-values compare cohorts using ANOVA or Kruskal–Wallis tests for continuous variables and χ2 test or Fisher’s exact test for categorical variables, as appropriate.
Models
AUC
NPV
PPV
Mayo-Internal
MIMIC-External
Mayo ED-External
Mayo-Internal
MIMIC-External
Mayo ED-External
Mayo-Internal
MIMIC-External
Mayo ED-External
Single modality without reasoning
Image only (CheXpert-DenseNet)
0.500 ± 0.004
0.526 ± 0.003
0.499 ± 0.007
0.684 ± 0.005
0.798 ± 0.003
0.883 ± 0.003
0.322 ± 0.005
0.234 ± 0.003
0.100 ± 0.002
Image only (MedCLIP-ViT)
0.539 ± 0.003
0.520 ± 0.003
0.489 ± 0.013
0.712 ± 0.004
0.789 ± 0.004
0.932 ± 0.003
0.347 ± 0.005
0.225 ± 0.003
0.051 ± 0.002
Text only (BioGPT)
0.740 ± 0.003
0.772 ± 0.004
0.808 ± 0.007
0.797 ± 0.003
0.893 ± 0.003
0.961 ± 0.002
0.518 ± 0.004
0.416 ± 0.005
0.315 ± 0.006
Text only (MedCLIP-ClinicalBERT)
0.743 ± 0.003
0.798 ± 0.004
0.804 ± 0.005
0.803 ± 0.004
0.903 ± 0.002
0.960 ± 0.002
0.515 ± 0.004
0.425 ± 0.006
0.294 ± 0.006
TABLE II : Performance comparison of single modality and multimodal models across internal and external datasets. Values are reported as point estimate ± half-width of the 95% confidence interval.
Fig. 2 : Subgroup analysis on External MIMIC datasets comparing Three models across AUC (above) and average green score (below).
Evaluation Metric
Mayo Internal
MIMIC External
Mayo ED External
Baseline
Proposed
P-value
Baseline
Proposed
P-value
Baseline
Proposed
P-value
Reasoning Quality
Factual correctness
0.405 ± 0.453
0.547 ± 0.425
p<0.001
0.692 ± 0.341
0.670 ± 0.385
p=0.404
0.269 ± 0.406
0.660 ± 0.392
p<0.001
Reasoning completeness
0.194 ± 0.297
0.321 ± 0.312
p<0.001
0.368 ± 0.260
0.285 ± 0.227
p<0.001
0.091 ± 0.202
0.364 ± 0.285
p<0.001
Hallucination ↓
0.587 ± 0.456
0.447 ± 0.424
p<0.001
0.302 ± 0.341
0.326 ± 0.383
p=0.329
0.724 ± 0.411
0.337 ± 0.390
p<0.001
Clinical Utility
TABLE III : Evaluation of multimodal clinical reasoning quality and utility. Reasoning quality metrics are reported as mean ± standard deviation, while clinical utility metrics are reported as point estimate ± half-width of the 95% confidence interval. ‘–’ means not applicable.
Fig. 3 : Reasoning quality evaluation. (a) Example true-negative case comparing baseline and proposed model reasoning. The baseline model incorrectly predicted high risk without considering relevant CXR findings, whereas the proposed RL-based model correctly predicted low risk with clinically grounded reasoning. (b) Expert preference comparison between baseline and proposed model reasoning across user groups.
Ablation
Mayo-Internal
MIMIC-External
Mayo ED-External
Lingshu-7B
AUC
NPV
PPV
Green Score
AUC
NPV
PPV
Green Score
AUC
NPV
PPV
Green Score
Modality contribution
Image-only
0.682 ± 0.006
0.784 ± 0.004
0.471 ± 0.006
–
0.638 ± 0.004
0.852 ± 0.003
0.305 ± 0.004
–
0.750 ± 0.007
0.949 ± 0.003
0.206 ± 0.006
–
Text-only
0.693 ± 0.003
0.788 ± 0.004
0.466 ± 0.003
0.577 ± 0.416
0.762 ± 0.003
0.897 ± 0.001
0.387 ± 0.004
0.832 ± 0.282
0.752 ± 0.009
0.946 ± 0.002
0.238 ± 0.009
0.759 ± 0.310
Representation learning
No Text Adapter
0.639 ± 0.002
0.746 ± 0.003
0.437 ± 0.003
0.429 ± 0.364
0.719 ± 0.007
0.874 ± 0.004
0.367 ± 0.007
0.728 ± 0.296
0.733 ± 0.006
0.944 ± 0.002
0.208 ± 0.007
0.626 ± 0.309
TABLE IV : Ablation study of Lingshu-7B across internal and external datasets. ‘–’ means not applicable
Cardiovascular disease (CVD) remains the leading cause of death among women, yet cardiovascular risk assessment often relies on clinical variables that may be missing, outdated, or unavailable in routine care. Screening mammography offers an opportunity for opportunistic cardiovascular risk stratification because it is routinely acquired and contains vascular features, including breast arterial calcifications (BAC), that are associated with cardiovascular risk and events. We evaluate whether mammography specific foundation models, originally pretrained for breast cancer-related tasks, can transfer to cardiovascular risk prediction without cardiovascular specific supervision or explicit BAC annotation. We constructed a 5-year major adverse cardiovascular event (MACE) cohort of 22,497 women linked to electronic health record outcomes, including 500 events (2.22% prevalence). The foundation models achieved AUROCs of 0.823 and 0.822 substantially exceeding an age-only model (AUROC 0.765), despite using only the screening mammogram as input, with no clinical variables. Both foundation models evaluated assigned substantially higher predicted risk to patients with radiologist-documented BAC, despite BAC never being used as a training label, and showed activation patterns consistent with vascular findings. Together, these findings suggest that mammography foundation models can recover clinically relevant cardiovascular risk information directly from mammographic pixels and suggest that screening mammography may provide an opportunistic source of cardiovascular risk information to complement conventional clinical assessment without additional imaging. Code is available in https://github.com/PauFeld/MammoCVD
Paula Feldman, Nusrat Binta Nizam, Sunwoo Kwak +3
Department of Radiology, Weill Cornell Medicine, New York, NY, USA · Cornell Tech, New York, NY, USA · Cornell University, New York, NY, USA
Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA benchmarks focus on single images or short-horizon image pairs. We introduce MI-CXR, a benchmark for standardized evaluation of Multi-Interval longitudinal reasoning over multi-visit CXR sequences, without requiring free-form report generation or additional clinical context. MI-CXR comprises five-way multiple-choice questions over five-visit patient timelines and instantiates three complementary task families: Temporal Event Localization, Interval-wise Change Reasoning, and Global Trajectory Summarization, which assess clinically grounded visual reasoning over time. Evaluating 14 state-of-the-art vision-language models (VLMs) shows low overall performance, with an average accuracy of 29.3%, only modestly above random guessing. Using stage-wise diagnostic probing, we find that models often produce locally plausible interval descriptions but fail to enforce temporal constraints or compose evidence into globally consistent decisions over the full timeline. These findings reveal key limitations of current VLMs and establish MI-CXR as a principled benchmark for longitudinal medical reasoning. The benchmark is available at https://github.com/AIDASLab/MI-CXR
Sunghwan Steve Cho, Yunseok Han, Jaeyoung Do
AIDAS Laboratory, 1ECE & 2IPAI, Seoul National University
Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clinical trustworthiness. While visual attribution methods are widely used to explain LVLM predictions, whether these explanations actually reflect the visual evidence underlying the model's decision is largely unverified, since ground-truth annotations for internal model reasoning are typically unavailable. We address this question for chest X-ray (CXR) reasoning by developing a causal evaluation framework that retains only CXR-VQA samples for which the expert-annotated region is verified, via counterfactual editing, to be causally responsible for the model's prediction. Using this framework across 11 attribution methods, six open-source LVLMs, and two output modes (direct answer and step-by-step reasoning), we find that existing attribution methods often fail to identify the evidence used by LVLMs. To address this failure, we propose MedFocus, a concept-based attribution method that localizes clinically meaningful anatomical regions via unbalanced optimal transport and measures their causal effect on model outputs through targeted interventions. MedFocus produces spatial, concept-level, and token-level attributions and substantially outperforms prior methods, taking a step toward more trustworthy attribution for medical LVLMs. Our data and code are available at https://github.com/gzxiong/medfocus/.
Guangzhi Xiong, Qiao Jin, Sanchit Sinha +2
University of Virginia · National Institutes of Health