cs.CVSep 29, 2026

SAM Meets VLM: Parameter-Decoupled Full-Parameter Training for Unified Medical Reasoning and Segmentation

Authors: Xuyang Cao, Enyou Liu, Jun Zhao, Zhuoyun Liu, Jintao Fei, Leo

Organizations: JDH Algo, JD Health International Inc.

Abstract

Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special <SEG> token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the <SEG> hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable <SEG> prompt states, monitored by the Davies--Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

    Aug 11, 2026Yuan Wang, Hualiang Wang, Yixin Chen +6Medical Vision-Language ModelsPerception

  2. MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

    Jan 14, 2026Yang Xing, Jiong Wu, Savas Ozdemir +43D Medical Image SegmentationMedical Vision-Language Models

  3. MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models

    Jun 4, 2026Aofei Chang, Le Huang, Alex James Boyd +4Medical Vision-Language ModelsWeak Visual Grounding