SAM Meets VLM: Parameter-Decoupled Full-Parameter Training for Unified Medical Reasoning and Segmentation
Organizations: JDH Algo, JD Health International Inc.
Abstract
Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special <SEG> token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the <SEG> hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable <SEG> prompt states, monitored by the Davies--Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.
Figures & tables
| Model | MeCOVQA-G+ seg | MedSAM2 det | |||||||
|---|---|---|---|---|---|---|---|---|---|
| DER | CT | PET | X-RAY | END | MR | US | FP | ||
| MedPLIB 14B [ 19 ] | 79.84 | 57.58 | 64.25 | 8.47 ⋆ | 44.35 ⋆ | 27.38 ⋆ | 34.22 ⋆ | 4.82 ⋆ | – |
| UniBioMed † [ 82 ] | 63.29 | 14.07 | 19.03 | 29.86 | 67.35 | 25.74 | 0.18 | 69.22 | – |
| Qwen2.5-VL-7B+SAM2 † [ 2 , 59 ] | 32.97 | 19.31 | 13.98 | 5.75 | 52.91 | 7.79 | 8.48 | 0.32 | 20.90 |
| Qwen2.5-VL-32B+SAM2 † [ 2 , 59 ] | 34.94 | 23.23 | 22.84 | 4.81 | 58.55 | 10.75 | 10.55 | 0.22 | 30.10 |
| Ours-8B | 92.09 | 64.04 | 77.93 | 14.69 | 92.80 | 43.07 | 83.83 | 74.07 | 44.60 |
| Model | ISIC16 | Kvasir | IDRiD | CovidQUEx | Promise12 | MosMed+ | US-Nerve | TNBC |
|---|---|---|---|---|---|---|---|---|
| Unet-MIT | 89.10 | 56.90 | 5.30 | – | – | 76.10 | – | 75.90 |
| Unet-EfficientNet | 90.30 | 81.20 | 7.80 | 74.40 | 89.20 | 78.10 | 78.70 | 73.80 |
| Unet-MobileNetV2 | 89.10 | 75.40 | 9.20 | 74.20 | 89.60 | 78.50 | 77.20 | 76.20 |
| Unet-DenseNet121 | 89.30 | 79.40 | 8.90 | 75.60 | 90.00 | 79.10 | 78.60 | 78.80 |
| Unet-ResNet50 | 88.70 | 69.80 | 9.00 | 73.40 | 88.80 | 79.00 | 77.60 | 78.50 |
| Ours-8B | 93.16 | 90.91 | 47.10 | 77.84 | 90.81 | 78.17 | 80.15 | 82.93 |
| Model | VQA-RAD | MedXpertQA | SLAKE | PATH-VQA | PMC-VQA | Avg. |
|---|---|---|---|---|---|---|
| Qwen2.5-VL 7B [ 2 ] | 66.30 | 20.75 | 67.86 | 42.30 | 50.86 | 49.61 |
| Lingshu 7B [ 83 ] | 68.74 | 26.90 | 82.90 | 60.23 | 55.77 | 58.91 |
| HealthGPT 14B [ 36 ] | 64.08 | 24.55 | 67.43 | 58.67 | 56.90 | 54.33 |
| MedGemma 27B [ 65 ] | 63.86 | 33.10 | 76.17 | 47.60 | 45.35 | 53.22 |
| Qwen2.5-VL 32B [ 2 ] | 72.28 | 25.30 | 76.36 | 41.58 | 53.58 | 53.82 |
| Lingshu 32B [ 83 ] | 75.39 | 31.00 | 87.68 | 64.76 | 57.23 | 63.21 |
| Model | PubMedQA | MedMCQA | MedQA | MedXpertQA | CMMLU | Avg. |
|---|---|---|---|---|---|---|
| Qwen2.5-VL 7B [ 2 ] | 75.80 | 53.40 | 57.50 | 12.40 | 68.80 | 53.58 |
| Lingshu 7B [ 83 ] | 75.40 | 56.13 | 63.39 | 16.45 | 69.02 | 56.08 |
| HealthGPT 14B [ 36 ] | 69.40 | 63.33 | 66.93 | 12.45 | 55.36 | 53.49 |
| MedGemma 27B [ 65 ] | 79.00 | 63.23 | 81.15 | 22.01 | 60.24 | 61.13 |
| Qwen2.5-VL 32B [ 2 ] | 68.60 | 62.71 | 71.33 | 15.88 | 82.60 | 60.22 |
| Lingshu 32B [ 83 ] | 78.20 | 65.05 | 74.94 | 22.86 | 82.37 | 64.69 |
| Model | VQA | Text QA | Dice |
|---|---|---|---|
| w/o two-phase | 56.67 | 49.27 | 64.92 |
| w/ two-phase | 58.39 | 56.58 | 80.13 |
| Metric | ||||
|---|---|---|---|---|
| VQA Avg. | 56.23 | 56.90 | 57.93 | 58.84 |
| Text Avg. | 54.22 | 54.92 | 55.60 | 56.13 |
| Dice Avg. | 65.95 | 62.39 | 57.42 | 50.38 |
| Model | ISIC16 | Kvasir | IDRiD | CovidQUEx | Promise12 | MosMed+ | US-Nerve | TNBC | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| w/o Stage 3 | 86.19 | 83.10 | 41.34 | 72.49 | 89.99 | 79.68 | 80.10 | 77.01 | 76.24 |
| w/ Stage 3 | 93.71 | 91.22 | 56.86 | 78.39 | 91.40 | 78.01 | 80.24 | 81.20 | 81.38 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Included Public Datasets |
|---|---|
| Image–caption | PMCOA [ 37 ] , ROCO [ 53 ] , LLaVA-Med [ 32 ] , MedPix2.0 [ 70 ] , CheXpert Plus [ 5 ] , MIMIC-CXR [ 25 ] , ROCOv2 [ 62 ] , Quilt-LLaVA [ 67 ] , PubMedVision [ 9 ] , IU-Xray [ 17 ] |
| MM instruction | VQA-RAD [ 31 ] , PMC-VQA [ 88 ] , PATH-VQA [ 47 ] , SLAKE [ 39 ] , MIMIC-CXR-VQA [ 1 ] , VQA-Med-2019 [ 4 ] |
| Referring segmentation / grounding | SA-Med2D-20M [ 84 ] |
| Textual QA | JMed [ 75 ] , HealthCareMagic [ 35 ] , iCliniq [ 35 ] , HuatuoGPT2-SFT-GPT4 [ 10 ] , Citrus-S3 [ 75 ] , medical-o1-verifiable-problem [ 7 ] , Medical-R1-Distill-Data [ 7 ] , huatuogpt-o1-for-reasoning [ 7 ] , MedReason [ 81 ] , MedThoughts [ 20 ] , medical-o1-reasoningSFT [ 7 ] , AlpaCare [ 89 ] , ApolloCorpus [ 77 ] , MedQuAD [ 3 ] , MedQA [ 23 ] , PMC-LLaMA [ 78 ] |
| General instruction | LLaVA1.5 [ 40 ] , PixMo [ 16 ] , ALLaVA [ 6 ] , OpenHermes-2.5 [ 73 ] , OKVQA [ 44 ] , A-OKVQA [ 64 ] , OCRVQA [ 45 ] , TextCaps [ 69 ] |
| Axis | Benchmarks | Modality | Metric |
|---|---|---|---|
| Medical textual QA | PubMedQA [ 24 ] , MedMCQA [ 51 ] , MedQA [ 23 ] , MedXpertQA [ 91 ] , CMMLU [ 34 ] | Text | Accuracy |
| Medical visual QA | VQA-RAD [ 31 ] , MedXpertQA [ 91 ] , SLAKE [ 39 ] , PATH-VQA [ 47 ] , PMC-VQA [ 88 ] | Image+Text | Accuracy |
| Referring segmentation | MedSegBench [ 29 ] , MeCOVQA-G [ 76 ] | Image+Text | Dice |
| Grounding | MedSAM2 det [ 76 ] | Image+Text | Precision@0.5 |
| Subset | Domain / Modality | Segmentation Target | Metric |
|---|---|---|---|
| ISIC16 [ 13 ] | Dermoscopy | Skin lesion region in dermoscopic images | Dice |
| Kvasir [ 22 ] | Endoscopy | Colorectal polyp region in colonoscopy images | Dice |
| IDRiD [ 55 ] | Fundus photography | Retinal pathology or structure masks for diabetic retinopathy analysis | Dice |
| CovidQUEx [ 71 ] | Chest X-ray | Lung or infection-related regions in COVID-19 radiographs | Dice |
| Promise12 [ 38 ] | Prostate MRI | Prostate gland region in pelvic MR images | Dice |
| MosMed+ [ 46 ] | Chest CT | Lung or infection-related regions in thoracic CT volumes | Dice |