cs.CVSep 28, 2026

Clinical Trajectory Alignment for Medical Vision-Language Pre-training

Authors: Huimin Yan, Xian Yang, Zhi Wang, Liang Bai

Organizations: Institute of Intelligent Information Processing, Shanxi University, Taiyuan, China · Alliance Manchester Business School, The University of Manchester, Manchester, UK

Abstract

Medical vision-language pre-training largely follows a visit-level image-report matching paradigm, aligning paired images and reports at individual visits. While effective for static cross-modal correspondence, this paradigm provides limited supervision for longitudinal clinical change, such as whether abnormalities improve, remain stable, or worsen over time. Learning such change is challenging because temporal semantics are implicit in free-text reports, and different abnormalities within the same patient may evolve asynchronously or even in opposite directions. We propose MedCTA, which reframes medical vision-language pre-training from visit-level cross-modal matching to learning clinical change. Rather than compressing a patient history into a single temporal representation, MedCTA models clinical change at two complementary scopes. At the abnormality scope, clinically grounded queries construct abnormality-conditioned visual and textual trajectories to capture heterogeneous abnormality evolution. At the patient-course scope, global image and report sequences are modeled to capture overall clinical progression beyond any individual abnormality. Structured trend supervision is extracted from longitudinal reports by an offline LLM parser, removing the need for manual temporal annotations. Combined with static image-report alignment, MedCTA learns representations that preserve visit-level cross-modal correspondence while encoding longitudinal change semantics. Experiments on temporal image classification, image-text retrieval, and zero-shot classification show consistent gains over strong medical vision-language baselines.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CT-ΔΔBench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

    Aug 12, 2026Kegeng Tang, Jingbo Wang, Shaogang Ren +1Longitudinal ImagingReport

  2. LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

    Jul 30, 2026Zhilin Wu, Zhangkai Ni, Chengmei Yang +4Medical Visual Question AnsweringMultimodal Clinical Data

  3. Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

    Jun 24, 2026Bowen Shi, Weiwei Cao, Ruifeng Yuan +5Medical Vision-Language ModelsRadiology Report Generation