cs.CVOct 7, 2026

ΔΔRepresentation: Geometry Supervised Representation Learning of Phenotypes via Counterfactual Reasoning for Medical VLMs

Authors: Hao Wang, Qiwei Zeng, Jinghao Lin, Shuchang Ye, Yuezhe Yang, Yige Peng, Haoyuan Che, Jinman Kim, +1 more

Organizations: The University of Sydney · Shanghai Jiao Tong University · Northeastern University · Jilin University

Abstract

Medical vision-language models (VLMs) have shown increasing potential for radiological image interpretation. Medical VLMs encode radiological images into visual representations that capture both anatomical and phenotypic information for diagnosis. Existing approaches improve pathological phenotype representations through semantic-guided representation alignment. However, pathological phenotypes arise as lesion-specific visual changes superimposed on underlying normal anatomy. Such semantic alignment approaches fail to model the phenotype-specific increment relative to the corresponding normal anatomical representation. To address this gap, we propose \textbf{ΔΔRepresentation}, a visual phenotype representation learning framework based on counterfactual reasoning for medical VLMs. It comprises \textbf{BaseAnatomy}, a geometry-supervised representation learning module, and \textbf{ΔΔPhenotype}, a counterfactual incremental representation learning module. BaseAnatomy provides fine-grained geometric supervision through spatial relationships across and within anatomical structures. ΔΔPhenotype computes the representation increment between lesion representations and their corresponding normal anatomical representations, and supervises increments associated with the same phenotype to cluster in the representation space. Experiments on \textit{ReXGroundingCT} and \textit{LIDC-IDRI} demonstrate that ΔΔRepresentation effectively structures pathological phenotype representations and improves lesion grounding and phenotype characterization accuracy in medical VLMs. Code is available at https://anonymous.4open.science/r/deltarep-CF6D.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Geometry-Supervised Visual Representation Learning for Multi-Phenotype Lesion Interpretation in Medical VLMs

    Oct 7, 2026Hao Wang, Qiwei Zeng, Shuchang Ye +5Medical Vision-Language Models

  2. EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

    Jun 4, 2026Hao Wang, Qiwei Zeng, Jinghao Lin +6Medical Vision-Language ModelsPathology Foundation Models

  3. Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

    Jul 4, 2026Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj +2Medical Vision-Language ModelsMedical Visual Question Answering