Clinical Trajectory Alignment for Medical Vision-Language Pre-training
Authors: Huimin Yan, Xian Yang, Zhi Wang, Liang Bai
Organizations: Institute of Intelligent Information Processing, Shanxi University, Taiyuan, China · Alliance Manchester Business School, The University of Manchester, Manchester, UK
Medical vision-language pre-training largely follows a visit-level image-report matching paradigm, aligning paired images and reports at individual visits. While effective for static cross-modal correspondence, this paradigm provides limited supervision for longitudinal clinical change, such as whether abnormalities improve, remain stable, or worsen over time. Learning such change is challenging because temporal semantics are implicit in free-text reports, and different abnormalities within the same patient may evolve asynchronously or even in opposite directions. We propose MedCTA, which reframes medical vision-language pre-training from visit-level cross-modal matching to learning clinical change. Rather than compressing a patient history into a single temporal representation, MedCTA models clinical change at two complementary scopes. At the abnormality scope, clinically grounded queries construct abnormality-conditioned visual and textual trajectories to capture heterogeneous abnormality evolution. At the patient-course scope, global image and report sequences are modeled to capture overall clinical progression beyond any individual abnormality. Structured trend supervision is extracted from longitudinal reports by an offline LLM parser, removing the need for manual temporal annotations. Combined with static image-report alignment, MedCTA learns representations that preserve visit-level cross-modal correspondence while encoding longitudinal change semantics. Experiments on temporal image classification, image-text retrieval, and zero-shot classification show consistent gains over strong medical vision-language baselines.
Figures & tables
Figure 1: Motivation of MedCTA. (a) Static image–report matching captures visit-level correspondence but overlooks longitudinal clinical change. (b) MedCTA exploits follow-up records with LLM-derived trend supervision to model abnormality-specific and patient-course trajectories.
Figure 2: Overview of MedCTA. Given a longitudinal image–report sequence Si , MedCTA jointly performs static image–report alignment, Abnormality-Specific Trajectory Alignment, and Patient-Course Progression Alignment. Abnormality-specific and patient-course trajectories are temporally encoded by Mamba and constrained by LLM-derived trend supervision, enabling semantic-level trajectory alignment for abnormality evolution and overall clinical progression.
Method
Consolidation
Edema
Pl. Effusion
Pneumonia
Pneumothorax
MGCA
44.99 ± 0.47
62.71 ± 0.24
57.17 ± 0.69
63.73 ± 0.89
52.16 ± 1.02
MedCLIP
53.40 ± 0.87
42.85 ± 0.01
49.96 ± 0.28
67.10 ± 0.22
55.46 ± 0.02
PRIOR
41.85 ± 1.51
43.97 ± 1.03
59.21 ± 1.19
41.41 ± 3.24
45.33 ± 2.13
MAVL
43.79 ± 0.00
42.85 ± 0.00
48.66 ± 0.00
67.10 ± 0.00
55.46 ± 0.00
CARZero
48.95 ± 1.25
51.92 ± 1.64
50.19 ± 1.72
57.44 ± 1.35
47.84 ± 1.78
Med-ST
60.57 ± 1.18
67.35 ± 0.32
58.47 ± 1.50
65.00 ± 0.34
54.18 ± 0.81
Table 1: Temporal image classification results. Accuracy (%) is reported across five abnormality categories. Best results are shown in bold, and second-best results are underlined.
Figure 3: Accuracy (%) for temporal sentence similarity classification.
Text → Image
Image → Text
Method
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
MGCA
74.50
76.30
71.16
60.85
63.62
63.88
64.11
61.95
MedCLIP
45.75
47.10
48.63
43.07
50.31
48.37
48.30
48.09
PRIOR
47.13
48.11
47.53
47.24
49.50
52.55
51.95
36.55
MAVL
50.00
50.00
42.91
37.27
50.00
50.00
49.47
50.00
CARZero
52.44
50.00
49.47
50.01
50.00
50.00
51.65
47.38
Table 2: Cross-modal retrieval results on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision (%) at top- k (P@ k ). Best results are shown in bold, and second-best results are underlined.
Figure 4: Qualitative dynamic phrase grounding results on Chest ImaGenome for improving, stable, and worsening cases. The left and right images denote the prior and current studies, respectively, with highlighted regions indicating visual evidence related to temporal progression.
Figure 5: Qualitative examples of image-to-text and text-to-image retrieval on MIMIC-5 × 200. For each method, we show the top two retrieved samples given an input chest X-ray or a query sentence. Green and red boxes denote correct and incorrect retrievals, respectively, with abnormality categories annotated below the incorrect samples.
RSNA Pneumonia
COVIDx
Method
ACC
F1
ACC
F1
MGCA
85.06 ± 0.05
76.56 ± 0.10
92.13 ± 0.10
80.10 ± 0.65
MedCLIP
84.17 ± 0.16
77.38 ± 0.35
90.19 ± 0.09
75.60 ± 2.10
PRIOR
81.30 ± 0.35
72.02 ± 1.09
90.01 ± 0.19
87.97 ± 0.68
MAVL
79.99 ± 0.83
73.14 ± 0.48
88.77 ± 0.92
85.09 ± 1.71
CARZero
84.28 ± 0.07
76.94 ± 0.40
91.92 ± 0.20
85.83 ± 0.98
Table 3: Zero-shot classification results on the RSNA Pneumonia and COVIDx datasets. Performance is reported in terms of Accuracy (ACC, %) and F1 (%). Best results are shown in bold, and second-best results are underlined.
Method
Temporal Sentence Similarity
Temporal Image Classification
Consolidation
Edema
Pleural Effusion
Pneumonia
Pneumothorax
Ours
95.72
67.14
70.30
65.59
72.85
61.66
w/o clip loss
92.24
59.58
59.51
63.44
67.50
56.75
( − 3.48)
( − 7.56)
( − 10.79)
( − 2.15)
( − 5.35)
( − 4.91)
w/o abnormality loss
87.47
55.22
51.46
58.41
65.24
53.33
( − 8.25)
( − 11.92)
( − 18.84)
( − 7.18)
( − 7.61)
( − 8.33)
Table 4: Ablation study on temporal tasks on the MS-CXR-T dataset. Accuracy (%) for temporal sentence similarity and temporal image classification. Relative changes from the full model are shown below.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Text → Image
Image → Text
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
MedCTA
80.24
79.08
75.68
65.82
69.13
68.63
67.23
63.01
w/o Lclip
58.05
55.35
54.32
46.06
53.22
53.19
53.98
48.72
( − 22.19)
( − 23.73)
( − 21.36)
( − 19.76)
( − 15.91)
( − 15.44)
( − 13.25)
( − 14.29)
w/o Labn
77.67
76.59
73.45
63.93
67.51
66.36
64.75
61.28
( − 2.57)
( − 2.49)
( − 2.23)
( − 1.89)
( − 1.62)
( − 2.27)
( − 2.48)
( − 1.73)
Appendix
Table 5: Additional ablation results on image–text retrieval on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision (%) at top- k (P@ k ). Relative changes from the full model are shown below.
Method
Text → Image
Image → Text
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
Random init
78.37
77.69
72.94
65.32
68.27
65.42
68.15
61.82
Clinical init
80.24
79.08
75.68
65.82
69.13
68.63
67.23
63.01
Appendix
Table 6: Effect of abnormality query initialization on image–text retrieval on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision at top- k (P@k).
Method
Text → Image
Image → Text
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
Qwen2-7B
80.24
79.08
75.68
65.82
69.13
68.63
67.23
63.01
Llama 3-8B
80.11
78.92
76.89
66.09
69.56
68.54
66.85
63.14
Appendix
Table 7: Robustness to different LLM parsers on image–text retrieval on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision at top- k (P@k).
In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrence detection, and ongoing patient management. Yet, despite this central role of temporal comparison in clinical decision-making, existing medical foundation models remain largely confined to single-study understanding, leaving temporally grounded cross-examination insufficiently addressed. To address this gap, we study longitudinal imaging difference reporting, a task in which a model takes two temporally separated scans from the same patient and generates a clinically meaningful report describing interval changes between them. We introduce CT-ΔBench, a dedicated benchmark for this task with patient-level splitting to prevent information leakage. To better evaluate this task beyond surface-level text similarity, we further develop change-aware metrics specifically designed to capture clinically meaningful longitudinal changes, and conduct an independent physician validation to assess the reliability of the synthesized references and event extraction pipeline. We also compare direct paired-CT reasoning with an indirect two-stage pipeline that first generates single-timepoint reports and then performs textual differencing. Finally, we propose DeltaMed, a baseline model for direct paired-CT difference reporting, and train it on the benchmark training set. Together, these contributions lay the groundwork for temporally aware medical foundation models that better reflect real-world longitudinal clinical reasoning.
In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response. However, despite the rapid advancement of multimodal large language models (MLLMs), longitudinal medical visual reasoning remains largely underexplored. To fill this gap, we propose LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis. LoMeVQA covers five tasks: progress classification, progress description, progress report generation, differential region grounding, and differential region description. To construct the dataset, we develop an automated pipeline that (1) organizes patient records chronologically, (2) extracts clinically meaningful entities via a medical knowledge graph, and (3) models their temporal evolution to guide large language models in generating high-quality longitudinal VQA pairs. Extensive evaluations demonstrate that both general-purpose and medical-domain MLLMs perform poorly on LoMeVQA, revealing substantial limitations in temporal reasoning. To address these limitations, we introduce MedLong-8B, which achieves state-of-the-art performance across all tasks. Beyond benchmarking, we conduct detailed analyses that uncover key failure modes and shed light on how to improve longitudinal medical visual reasoning. Our data is available at: https://github.com/pepperbubble/LoMeVQA
Zhilin Wu, Zhangkai Ni, Chengmei Yang +4
Tongji University Shanghai, China · East China Normal University Shanghai, China
Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones and coarse semantic alignment. To address these issues, we propose a tailored VLP framework featuring three key components: (1) a CNN-ViT hybrid encoder that replaces ViT's patch embedding with a 3D CNN backbone to efficiently capture local anatomical details while preserving global attention and compatibility with pre-trained cross-modal priors; (2) a disease-level contrastive learning mechanism using learnable query tokens to dynamically extract disease-specific semantics from full reports and align them with corresponding visual features, thereby disentangling distinct diseases within the same anatomical region; and (3) a diagnosis-aware prompt strategy that employs real clinical phrases and aggregated disease prototypes to bridge the pre-training-inference gap and enhance zero-shot diagnostic reliability. Our model achieves state-of-the-art performance on CT-RATE (84.4% AUC, +5.1%) and Rad-ChestCT (75.4% AUC, +5.4%), with even larger gains (+9.8% AUC) on a challenging 60-disease benchmark, and demonstrates strong transferability to radiology report generation, underscoring the generality and clinical utility of our approach.
Bowen Shi, Weiwei Cao, Ruifeng Yuan +5
DAMO Academy, Alibaba Group · Hupan Lab, 310023, Hangzhou, China · Shanghai Jiao Tong University, China +2