Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special <SEG> token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the <SEG> hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable <SEG> prompt states, monitored by the Davies--Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.
Figures & tables
Figure 1: Unified medical MLLM architecture and parameter-decoupled training overview. A VLM produces language responses for reasoning tasks and emits a <SEG> token for language-conditioned referring segmentation. The <SEG> hidden state is projected as an implicit prompt to the SAM2-based segmentation branch, while the bottom row summarizes the three training stages.
Figure 2: t-SNE [ 74 ] visualization of <SEG> token hidden states as training progresses. DBI decreases from 3.75 to 1.75 across the two displayed checkpoints, showing that the implicit segmentation prompts become more compact and separable across target classes.
Model
MeCOVQA-G+ seg
MedSAM2 det
DER
CT
PET
X-RAY
END
MR
US
FP
MedPLIB 14B [ 19 ]
79.84
57.58
64.25
8.47 ⋆
44.35 ⋆
27.38 ⋆
34.22 ⋆
4.82 ⋆
–
UniBioMed † [ 82 ]
63.29
14.07
19.03
29.86
67.35
25.74
0.18
69.22
–
Qwen2.5-VL-7B+SAM2 † [ 2 , 59 ]
32.97
19.31
13.98
5.75
52.91
7.79
8.48
0.32
20.90
Qwen2.5-VL-32B+SAM2 † [ 2 , 59 ]
34.94
23.23
22.84
4.81
58.55
10.75
10.55
0.22
30.10
Ours-8B
92.09
64.04
77.93
14.69
92.80
43.07
83.83
74.07
44.60
Table 1: Results on multi-modal referring segmentation (Dice, %) and medical detection (precision@0.5). Bold: best; underline: second best; –: not supported; ⋆ : zero-shot; † : reproduced/evaluation-only baseline.
Model
ISIC16
Kvasir
IDRiD
CovidQUEx
Promise12
MosMed+
US-Nerve
TNBC
Unet-MIT
89.10
56.90
5.30
–
–
76.10
–
75.90
Unet-EfficientNet
90.30
81.20
7.80
74.40
89.20
78.10
78.70
73.80
Unet-MobileNetV2
89.10
75.40
9.20
74.20
89.60
78.50
77.20
76.20
Unet-DenseNet121
89.30
79.40
8.90
75.60
90.00
79.10
78.60
78.80
Unet-ResNet50
88.70
69.80
9.00
73.40
88.80
79.00
77.60
78.50
Ours-8B
93.16
90.91
47.10
77.84
90.81
78.17
80.15
82.93
Table 2: Segmentation results on MedSegBench [ 29 ] (Dice, %). Bold: best; underline: second best; –: not reported.
Model
VQA-RAD
MedXpertQA
SLAKE
PATH-VQA
PMC-VQA
Avg.
Qwen2.5-VL 7B [ 2 ]
66.30
20.75
67.86
42.30
50.86
49.61
Lingshu 7B [ 83 ]
68.74
26.90
82.90
60.23
55.77
58.91
HealthGPT 14B [ 36 ]
64.08
24.55
67.43
58.67
56.90
54.33
MedGemma 27B [ 65 ]
63.86
33.10
76.17
47.60
45.35
53.22
Qwen2.5-VL 32B [ 2 ]
72.28
25.30
76.36
41.58
53.58
53.82
Lingshu 32B [ 83 ]
75.39
31.00
87.68
64.76
57.23
63.21
Table 3: Medical visual question answering results (accuracy, %). Bold: best; underline: second best.
Model
PubMedQA
MedMCQA
MedQA
MedXpertQA
CMMLU
Avg.
Qwen2.5-VL 7B [ 2 ]
75.80
53.40
57.50
12.40
68.80
53.58
Lingshu 7B [ 83 ]
75.40
56.13
63.39
16.45
69.02
56.08
HealthGPT 14B [ 36 ]
69.40
63.33
66.93
12.45
55.36
53.49
MedGemma 27B [ 65 ]
79.00
63.23
81.15
22.01
60.24
61.13
Qwen2.5-VL 32B [ 2 ]
68.60
62.71
71.33
15.88
82.60
60.22
Lingshu 32B [ 83 ]
78.20
65.05
74.94
22.86
82.37
64.69
Table 4: Medical textual question answering results (accuracy, %). Bold: best; underline: second best.
Model
VQA
Text QA
Dice
w/o two-phase
56.67
49.27
64.92
w/ two-phase
58.39
56.58
80.13
Table 5: Ablation on two-phase instruction fine-tuning (Ours-8B).
Figure 3: Qualitative comparison of medical referring segmentation results. We compare MedPLIB, UniBioMed, Qwen2.5-VL-32B+SAM2, and Ours-33B across representative imaging modalities. Green masks denote ground truth, all predicted segmentation masks are shown in blue, and yellow boxes indicate the grounding output used by the Qwen2.5-VL+SAM2 pipeline.
Metric
1.0
0.3
0.01
0.001
VQA Avg.
56.23
56.90
57.93
58.84
Text Avg.
54.22
54.92
55.60
56.13
Dice Avg.
65.95
62.39
57.42
50.38
Table 6: Sensitivity to gradient scaling factor γseg (Ours-8B).
Department of Biomedical Engineering, University of Florida, Gainesville, FL, USA · Department of Radiology, University of Florida, Jacksonville, FL, USA · Research Computing, University of Florida, Gainesville, FL, USA +2
GE HealthCare · College of Information Sciences and Technology, The Pennsylvania State University, State College, PA, USA · GE Healthcare, Bellevue, WA, USA.