Medical vision-language models (VLMs) have shown increasing potential for radiological image interpretation. Medical VLMs encode radiological images into visual representations that capture both anatomical and phenotypic information for diagnosis. Existing approaches improve pathological phenotype representations through semantic-guided representation alignment. However, pathological phenotypes arise as lesion-specific visual changes superimposed on underlying normal anatomy. Such semantic alignment approaches fail to model the phenotype-specific increment relative to the corresponding normal anatomical representation. To address this gap, we propose \textbf{ΔRepresentation}, a visual phenotype representation learning framework based on counterfactual reasoning for medical VLMs. It comprises \textbf{BaseAnatomy}, a geometry-supervised representation learning module, and \textbf{ΔPhenotype}, a counterfactual incremental representation learning module. BaseAnatomy provides fine-grained geometric supervision through spatial relationships across and within anatomical structures. ΔPhenotype computes the representation increment between lesion representations and their corresponding normal anatomical representations, and supervises increments associated with the same phenotype to cluster in the representation space. Experiments on \textit{ReXGroundingCT} and \textit{LIDC-IDRI} demonstrate that ΔRepresentation effectively structures pathological phenotype representations and improves lesion grounding and phenotype characterization accuracy in medical VLMs. Code is available at https://anonymous.4open.science/r/deltarep-CF6D.
Figures & tables
Figure 1: BaseAnatomy structures the anatomical visual representation space, while Δ Phenotype organizes lesion-specific increments relative to counterfactual normal anatomy.
Figure 2: Overview of Δ Representation. (a) BaseAnatomy shapes an anatomically structured visual space with ideal geometric supervision that encodes anatomical relationships and spatial continuity. (b) Δ Phenotype constructs counterfactual normal references for lesion patches and learns phenotype-specific visual increments relative to normal anatomy for discriminative phenotype representation learning.
Models
Scale
ReXGroundingCT
LIDC-IDRI
Grounding Acc.
Phenotype Acc.
Grounding Acc.
Phenotype Acc.
RadFM
14B
10.60
4.95
18.18
2.48
w/ Δ Representation
∼ 14B
50.53 ↑ 39.93
38.87 ↑ 33.92
62.81 ↑ 44.63
47.11 ↑ 44.63
LLaVA-Med
7B
43.11
12.37
2.48
5.79
w/ Δ Representation
∼ 7B
66.78 ↑ 23.67
51.94 ↑ 39.57
47.11 ↑ 44.63
44.63 ↑ 38.84
Lingshu
7B
77.39
15.19
23.97
42.15
Table 1: Comparison with pretrained medical VLMs and alignment methods on the VQA task.
Figure 3: Anatomical representation analysis of BaseAnatomy. (A–B) Representation distributions before and after BaseAnatomy training. (C) Spatial continuity across anatomical positions. (D) Cross-scale consistency of anatomical representations.
Figure 4: Counterfactual representation analysis. Comparison between inferred counterfactual representations and their corresponding normal representations.
Lesion Data Scale
ReXGroundingCT (VQA)
ReXGroundingCT (RRG)
LIDC-IDRI (VQA)
LIDC-IDRI (RRG)
Grounding Acc.
Phenotype Acc.
Grounding Acc.
Phenotype Acc.
Grounding Acc.
Phenotype Acc.
Grounding Acc.
Phenotype Acc.
1%
54.06
25.80
10.95
15.55
43.80
54.55
40.50
50.41
5%
72.44 ↑ 18.38
34.63 ↑ 8.83
10.95 ↑ 0.00
15.55 ↑ 0.00
44.63 ↑ 0.83
62.81 ↑ 8.26
40.50 ↑ 0.00
51.24 ↑ 0.83
10%
75.27 ↑ 2.83
39.93 ↑ 5.30
22.26 ↑ 11.31
31.45 ↑ 15.90
60.33 ↑ 15.70
68.60 ↑ 5.79
43.80 ↑ 3.30
53.72 ↑ 2.48
25%
78.09 ↑ 2.82
40.64 ↑ 0.71
23.32 ↑ 1.06
43.11 ↑ 11.66
73.55 ↑ 13.22
77.69 ↑ 9.09
44.63 ↑ 0.83
69.42 ↑ 15.70
50%
79.15 ↑ 1.06
43.11 ↑ 2.47
28.27 ↑ 4.95
50.18 ↑ 7.07
76.03 ↑ 2.48
77.69 ↑ 0.00
47.11 ↑ 2.48
75.21 ↑ 5.79
Table 3: Performance on ReXGroundingCT and LIDC-IDRI with different proportions of lesion data for the VQA and RRG tasks.
Figure 5: Representation learning performance of visual phenotypes in different lesion data scales.
Figure 6: Case study of Δ Representation, illustrating anatomical projection by BaseAnatomy, phenotype-increment projection by Δ Phenotype, and their contribution to report generation.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Distribution of benchmark instances by dataset and task.
Models
Scale
Atelectasis Phenotype
Consolidation Phenotype
Ground-glass opacity Phenotype
Pulmonary nodule Phenotype
VQA Acc.
RRG Acc.
VQA Acc.
RRG Acc.
VQA Acc.
RRG Acc.
VQA Acc.
RRG Acc.
RadFM
14B
7.14
4.29
8.14
4.65
6.82
3.41
8.55
4.27
w/ Δ Representation
∼ 14B
44.29 ↑ 37.14
17.14 ↑ 12.86
48.84 ↑ 40.70
18.60 ↑ 13.95
36.93 ↑ 30.11
13.07 ↑ 9.66
49.15 ↑ 40.60
14.10 ↑ 9.83
LLaVA-Med
7B
28.57
24.29
29.07
29.07
26.70
18.75
27.78
23.93
w/ Δ Representation
∼ 7B
55.71 ↑ 27.14
45.71 ↑ 21.43
72.09 ↑ 43.02
45.35 ↑ 16.28
54.55 ↑ 27.84
30.68 ↑ 11.93
59.40 ↑ 31.62
39.74 ↑ 15.81
Lingshu
7B
51.43
1.43
56.98
2.33
40.34
2.27
45.30
2.56
Appendix
Table 4: Disease-wise VQA and RRG performance on ReXGroundingCT. Values are composite accuracies (%), obtained by pooling the Grounding and Phenotype decisions within each task. Gains are absolute percentage points over the native backbone.
(a) Absolute representation quality
Component
Evaluation cohort
Metric
Score
BaseAnatomy
ReX supervised pool
Rotation Pearson r↑
0.9742
Scaling Pearson r↑
0.9751
Translation Pearson r↑
0.9750
Δ Phenotypes
ReX four-phenotype set, n=379
Cosine silhouette ↑
0.3607
Nearest-prototype accuracy ↑
75.46%
Appendix
Table 5: High-dimensional representation diagnostics. Panel (a) reports absolute representation-quality measurements. Panel (b) isolates the effect of exact 4×4 spatial pooling. Here, Δ=4×−1× and pp denotes percentage points.
School of Computer Science, The University of Sydney, Sydney, NSW, Australia · Jilin University, Changchun, China · Northeastern University, Shenyang, China +1