Geometry-Supervised Visual Representation Learning for Multi-Phenotype Lesion Interpretation in Medical VLMs
Organizations: The University of Sydney · Shanghai Jiao Tong University · Northeastern University
Abstract
Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation. However, these models still struggle to interpret multi-phenotype lesions whose diagnosis requires the joint assessment of multiple pathological phenotypes. Existing vision-language alignment methods produce visual representations that fail to preserve anatomical hierarchies and relationships among phenotypic subclasses. This stems from their reliance on semantic supervision, which lacks geometric constraints to preserve these relationships in the visual embedding space. Moreover, the sparsity of lesion-related anatomical and phenotypic representations makes it difficult for medical VLMs to capture important diagnostic evidence. To address these limitations, we propose \textbf{PureVision}, a geometry-supervised visual representation learning framework for multi-phenotype lesion interpretation in medical VLMs. It combines a geometry-supervised representation learning module, \textbf{PureEyes}, and an anatomy-guided evidence aggregation module, \textbf{PureNeurons}. PureEyes provides geometric supervision through ideal spatial distributions that encode anatomical hierarchies and phenotypic subclass relationships. PureNeurons projects visual representations into the learned latent space, using their positions to selectively aggregate lesion-specific anatomical and phenotypic evidence. Experiments on \textit{LIDC-IDRI}, \textit{CBIS-DDSM}, and \textit{3DReasonKnee} demonstrate that PureVision improves lesion grounding and phenotype characterization in visual question answering and radiology report generation. Code is available at: https://anonymous.4open.science/r/purevision-06C2.
Figures & tables
| Models | Scale | LIDC-IDRI | CBIS-DDSM | 3DReasonKnee | |||
|---|---|---|---|---|---|---|---|
| Grounding Acc. | Phenotypes Acc. | Grounding Acc. | Phenotypes Acc. | Grounding Acc. | Phenotypes Acc. | ||
| RadFM | 14B | 2.00 | 5.00 | 2.00 | 17.00 | 5.00 | 3.00 |
| w/ PureVision | 14B | 11.00 +9.00 | 23.00 +18.00 | 23.00 +21.00 | 24.00 +7.00 | 21.00 +16.00 | 24.00 +21.00 |
| LLaVA-Med | 7B | 14.00 | 27.79 | 13.00 | 30.83 | 12.00 | 16.00 |
| w/ PureVision | 7B | 23.00 +9.00 | 44.86 +17.07 | 24.5 +11.50 | 37.50 +6.67 | 32.00 +20.00 | 31.00 +15.00 |
| Lingshu | 7B | 36.00 | 33.86 | 30.50 | 12.00 | 28.00 | 27.00 |
| Lesion Extraction Efficiency | Lesion Information Delivery Ablation | ||||
|---|---|---|---|---|---|
| Selected Patches ( ) | Lesion Patch Recall | Redundant Patch Rate | Method | Grounding Acc. | Phenotypes Acc. |
| 4 | 60.11 | 45.23 | MedGemma1.5 | 11.50 | 10.50 |
| 8 | 67.64 | 65.81 | w/o PureNeurons | 32.50 | 36.57 |
| 12 | 68.47 | 75.82 | w/ lesion grounding | 40.00 | 38.43 |
| 16 | 68.84 | 81.23 | w/ lesion grounding and soft semantic fusion | 45.50 | 49.93 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Models (VQA) | Density | Margin | Size | Sphericity | Calcification | Lobulation | Spiculation |
|---|---|---|---|---|---|---|---|
| RadFM | 4.00 | 3.00 | 3.00 | 5.50 | 7.00 | 8.50 | 4.00 |
| w/ PureVision | 11.00 +7.00 | 5.00 +2.00 | 5.00 +2.00 | 33.00 +27.50 | 40.00 +33.00 | 50.50 +42.00 | 16.50 +12.50 |
| LLaVA-Med | 17.50 | 7.50 | 5.00 | 39.00 | 46.50 | 56.50 | 22.50 |
| w/ PureVision | 20.50 +3.00 | 42.50 +35.00 | 27.00 +22.00 | 83.00 +44.00 | 60.50 +14.00 | 60.00 +3.50 | 20.50 -2.00 |
| Lingshu | 23.50 | 13.50 | 10.00 | 45.50 | 52.50 | 63.00 | 29.00 |
| w/ PureVision | 42.50 +19.00 | 27.50 +14.00 | 44.50 +34.50 | 30.50 -15.00 | 80.00 +27.50 | 60.00 -3.00 | 69.00 +40.00 |
| Models (VQA) | Mass Shape | Mass Margins | Calcification Distribution |
|---|---|---|---|
| RadFM | 11.00 | 18.50 | 21.50 |
| w/ PureVision | 33.50 +22.50 | 12.50 -6.00 | 26.00 +4.50 |
| LLaVA-Med | 33.50 | 27.50 | 31.50 |
| w/ PureVision | 33.00 -0.50 | 32.00 +4.50 | 47.50 +16.00 |
| Lingshu | 19.00 | 4.00 | 13.00 |
| w/ PureVision | 35.50 +16.50 | 18.50 +14.50 | 52.50 +39.50 |