Clinical Trajectory Alignment for Medical Vision-Language Pre-training
Authors: Huimin Yan, Xian Yang, Zhi Wang, Liang Bai
Organizations: Institute of Intelligent Information Processing, Shanxi University, Taiyuan, China · Alliance Manchester Business School, The University of Manchester, Manchester, UK
Medical vision-language pre-training largely follows a visit-level image-report matching paradigm, aligning paired images and reports at individual visits. While effective for static cross-modal correspondence, this paradigm provides limited supervision for longitudinal clinical change, such as whether abnormalities improve, remain stable, or worsen over time. Learning such change is challenging because temporal semantics are implicit in free-text reports, and different abnormalities within the same patient may evolve asynchronously or even in opposite directions. We propose MedCTA, which reframes medical vision-language pre-training from visit-level cross-modal matching to learning clinical change. Rather than compressing a patient history into a single temporal representation, MedCTA models clinical change at two complementary scopes. At the abnormality scope, clinically grounded queries construct abnormality-conditioned visual and textual trajectories to capture heterogeneous abnormality evolution. At the patient-course scope, global image and report sequences are modeled to capture overall clinical progression beyond any individual abnormality. Structured trend supervision is extracted from longitudinal reports by an offline LLM parser, removing the need for manual temporal annotations. Combined with static image-report alignment, MedCTA learns representations that preserve visit-level cross-modal correspondence while encoding longitudinal change semantics. Experiments on temporal image classification, image-text retrieval, and zero-shot classification show consistent gains over strong medical vision-language baselines.
Figures & tables
Figure 1: Motivation of MedCTA. (a) Static image–report matching captures visit-level correspondence but overlooks longitudinal clinical change. (b) MedCTA exploits follow-up records with LLM-derived trend supervision to model abnormality-specific and patient-course trajectories.
Figure 2: Overview of MedCTA. Given a longitudinal image–report sequence Si , MedCTA jointly performs static image–report alignment, Abnormality-Specific Trajectory Alignment, and Patient-Course Progression Alignment. Abnormality-specific and patient-course trajectories are temporally encoded by Mamba and constrained by LLM-derived trend supervision, enabling semantic-level trajectory alignment for abnormality evolution and overall clinical progression.
Method
Consolidation
Edema
Pl. Effusion
Pneumonia
Pneumothorax
MGCA
44.99 ± 0.47
62.71 ± 0.24
57.17 ± 0.69
63.73 ± 0.89
52.16 ± 1.02
MedCLIP
53.40 ± 0.87
42.85 ± 0.01
49.96 ± 0.28
67.10 ± 0.22
55.46 ± 0.02
PRIOR
41.85 ± 1.51
43.97 ± 1.03
59.21 ± 1.19
41.41 ± 3.24
45.33 ± 2.13
MAVL
43.79 ± 0.00
42.85 ± 0.00
48.66 ± 0.00
67.10 ± 0.00
55.46 ± 0.00
CARZero
48.95 ± 1.25
51.92 ± 1.64
50.19 ± 1.72
57.44 ± 1.35
47.84 ± 1.78
Med-ST
60.57 ± 1.18
67.35 ± 0.32
58.47 ± 1.50
65.00 ± 0.34
54.18 ± 0.81
Table 1: Temporal image classification results. Accuracy (%) is reported across five abnormality categories. Best results are shown in bold, and second-best results are underlined.
Figure 3: Accuracy (%) for temporal sentence similarity classification.
Text → Image
Image → Text
Method
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
MGCA
74.50
76.30
71.16
60.85
63.62
63.88
64.11
61.95
MedCLIP
45.75
47.10
48.63
43.07
50.31
48.37
48.30
48.09
PRIOR
47.13
48.11
47.53
47.24
49.50
52.55
51.95
36.55
MAVL
50.00
50.00
42.91
37.27
50.00
50.00
49.47
50.00
CARZero
52.44
50.00
49.47
50.01
50.00
50.00
51.65
47.38
Table 2: Cross-modal retrieval results on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision (%) at top- k (P@ k ). Best results are shown in bold, and second-best results are underlined.
Figure 4: Qualitative dynamic phrase grounding results on Chest ImaGenome for improving, stable, and worsening cases. The left and right images denote the prior and current studies, respectively, with highlighted regions indicating visual evidence related to temporal progression.
Figure 5: Qualitative examples of image-to-text and text-to-image retrieval on MIMIC-5 × 200. For each method, we show the top two retrieved samples given an input chest X-ray or a query sentence. Green and red boxes denote correct and incorrect retrievals, respectively, with abnormality categories annotated below the incorrect samples.
RSNA Pneumonia
COVIDx
Method
ACC
F1
ACC
F1
MGCA
85.06 ± 0.05
76.56 ± 0.10
92.13 ± 0.10
80.10 ± 0.65
MedCLIP
84.17 ± 0.16
77.38 ± 0.35
90.19 ± 0.09
75.60 ± 2.10
PRIOR
81.30 ± 0.35
72.02 ± 1.09
90.01 ± 0.19
87.97 ± 0.68
MAVL
79.99 ± 0.83
73.14 ± 0.48
88.77 ± 0.92
85.09 ± 1.71
CARZero
84.28 ± 0.07
76.94 ± 0.40
91.92 ± 0.20
85.83 ± 0.98
Table 3: Zero-shot classification results on the RSNA Pneumonia and COVIDx datasets. Performance is reported in terms of Accuracy (ACC, %) and F1 (%). Best results are shown in bold, and second-best results are underlined.
Method
Temporal Sentence Similarity
Temporal Image Classification
Consolidation
Edema
Pleural Effusion
Pneumonia
Pneumothorax
Ours
95.72
67.14
70.30
65.59
72.85
61.66
w/o clip loss
92.24
59.58
59.51
63.44
67.50
56.75
( − 3.48)
( − 7.56)
( − 10.79)
( − 2.15)
( − 5.35)
( − 4.91)
w/o abnormality loss
87.47
55.22
51.46
58.41
65.24
53.33
( − 8.25)
( − 11.92)
( − 18.84)
( − 7.18)
( − 7.61)
( − 8.33)
Table 4: Ablation study on temporal tasks on the MS-CXR-T dataset. Accuracy (%) for temporal sentence similarity and temporal image classification. Relative changes from the full model are shown below.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Text → Image
Image → Text
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
MedCTA
80.24
79.08
75.68
65.82
69.13
68.63
67.23
63.01
w/o Lclip
58.05
55.35
54.32
46.06
53.22
53.19
53.98
48.72
( − 22.19)
( − 23.73)
( − 21.36)
( − 19.76)
( − 15.91)
( − 15.44)
( − 13.25)
( − 14.29)
w/o Labn
77.67
76.59
73.45
63.93
67.51
66.36
64.75
61.28
( − 2.57)
( − 2.49)
( − 2.23)
( − 1.89)
( − 1.62)
( − 2.27)
( − 2.48)
( − 1.73)
Appendix
Table 5: Additional ablation results on image–text retrieval on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision (%) at top- k (P@ k ). Relative changes from the full model are shown below.
Method
Text → Image
Image → Text
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
Random init
78.37
77.69
72.94
65.32
68.27
65.42
68.15
61.82
Clinical init
80.24
79.08
75.68
65.82
69.13
68.63
67.23
63.01
Appendix
Table 6: Effect of abnormality query initialization on image–text retrieval on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision at top- k (P@k).
Method
Text → Image
Image → Text
P@1
P@2
P@5
P@10
P@1
P@2
P@5
P@10
Qwen2-7B
80.24
79.08
75.68
65.82
69.13
68.63
67.23
63.01
Llama 3-8B
80.11
78.92
76.89
66.09
69.56
68.54
66.85
63.14
Appendix
Table 7: Robustness to different LLM parsers on image–text retrieval on the MIMIC-CXR 5 × 200 benchmark. Performance is reported as precision at top- k (P@k).