From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning
Organizations: Dept. of Radiology, Mayo Clinic, Phoenix, AZ, USA · School of Computing and Augmented Intelligence, Arizona State University, Tempe, USA · Dept. of Radiology, Mayo Clinic, Phoenix, USA · Mayo Clinic, Phoenix, USA · Dept. of Cardiology, Mayo Clinic, Phoenix, USA
Abstract
Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical data offers a scalable approach for identifying high-risk individuals before acute events occur. Although chest X-rays (CXRs) capture latent cardiovascular biomarkers and clinical histories provide complementary patient context, existing medical vision-language models are primarily optimized for radiology interpretation rather than prognostic reasoning. We propose a causal reinforcement learning framework for multimodal clinical reasoning that integrates CXRs and physician-authored clinical histories for opportunistic MACE prediction. The framework introduces (1) a role-decoupled dual-LLM architecture that separates reasoning from risk prediction, (2) a dual-action causal reinforcement learning policy for evidence selection and reasoning optimization, and (3) causal token pruning to learn compact multimodal representations. Evaluated on an internal cohort, an emergency department cohort, and the external MIMIC dataset, the proposed framework consistently outperformed unimodal baselines and state-of-the-art medical vision-language models, achieving AUROCs of 0.720, 0.760, and 0.845, respectively. It also substantially improved reasoning quality, achieving higher GREEN scores and higher expert preference while maintaining robust predictive performance across diverse patient populations.
Figures & tables
| Characteristic | Total | Mayo Internal | MIMIC External | Mayo ED External | P-value |
| Number of patients | 22,573 | 15,802 | 4,616 | 2,155 | |
| Demographics | |||||
| Age (years) | 63.14 ± 17.34 | 64.16 ± 18.01 | 61.84 ± 17.33 | 58.38 ± 9.65 | – |
| Female sex, n (%) | 9,611 (42.6%) | 6,454 (40.6%) | 2,343 (50.8%) | 832 (38.6%) | |
| Male sex, n (%) | 13,001 (57.6%) | 9,431 (59.3%) | 2,272 (49.2%) | 1,323 (61.4%) | |
| Race, n (%) |
| Models | AUC | NPV | PPV | |||||||
| Mayo-Internal | MIMIC-External | Mayo ED-External | Mayo-Internal | MIMIC-External | Mayo ED-External | Mayo-Internal | MIMIC-External | Mayo ED-External | ||
| Single modality without reasoning | ||||||||||
| Image only (CheXpert-DenseNet) | 0.500 0.004 | 0.526 0.003 | 0.499 0.007 | 0.684 0.005 | 0.798 0.003 | 0.883 0.003 | 0.322 0.005 | 0.234 0.003 | 0.100 0.002 | |
| Image only (MedCLIP-ViT) | 0.539 0.003 | 0.520 0.003 | 0.489 0.013 | 0.712 0.004 | 0.789 0.004 | 0.932 0.003 | 0.347 0.005 | 0.225 0.003 | 0.051 0.002 | |
| Text only (BioGPT) | 0.740 0.003 | 0.772 0.004 | 0.808 0.007 | 0.797 0.003 | 0.893 0.003 | 0.961 0.002 | 0.518 0.004 | 0.416 0.005 | 0.315 0.006 | |
| Text only (MedCLIP-ClinicalBERT) | 0.743 0.003 | 0.798 0.004 | 0.804 0.005 | 0.803 0.004 | 0.903 0.002 | 0.960 0.002 | 0.515 0.004 | 0.425 0.006 | 0.294 0.006 | |
| Evaluation Metric | Mayo Internal | MIMIC External | Mayo ED External | ||||||
| Baseline | Proposed | P-value | Baseline | Proposed | P-value | Baseline | Proposed | P-value | |
| Reasoning Quality | |||||||||
| Factual correctness | 0.405 0.453 | 0.547 0.425 | 0.692 0.341 | 0.670 0.385 | 0.269 0.406 | 0.660 0.392 | |||
| Reasoning completeness | 0.194 0.297 | 0.321 0.312 | 0.368 0.260 | 0.285 0.227 | 0.091 0.202 | 0.364 0.285 | |||
| Hallucination | 0.587 0.456 | 0.447 0.424 | 0.302 0.341 | 0.326 0.383 | 0.724 0.411 | 0.337 0.390 | |||
| Clinical Utility | |||||||||
| Ablation | Mayo-Internal | MIMIC-External | Mayo ED-External | |||||||||
| Lingshu-7B | AUC | NPV | PPV | Green Score | AUC | NPV | PPV | Green Score | AUC | NPV | PPV | Green Score |
| Modality contribution | ||||||||||||
| Image-only | 0.682 0.006 | 0.784 0.004 | 0.471 0.006 | – | 0.638 0.004 | 0.852 0.003 | 0.305 0.004 | – | 0.750 0.007 | 0.949 0.003 | 0.206 0.006 | – |
| Text-only | 0.693 0.003 | 0.788 0.004 | 0.466 0.003 | 0.577 0.416 | 0.762 0.003 | 0.897 0.001 | 0.387 0.004 | 0.832 0.282 | 0.752 0.009 | 0.946 0.002 | 0.238 0.009 | 0.759 0.310 |
| Representation learning | ||||||||||||
| No Text Adapter | 0.639 0.002 | 0.746 0.003 | 0.437 0.003 | 0.429 0.364 | 0.719 0.007 | 0.874 0.004 | 0.367 0.007 | 0.728 0.296 | 0.733 0.006 | 0.944 0.002 | 0.208 0.007 | 0.626 0.309 |