MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
Authors: Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
Organizations: Department of Biomedical Engineering, University of Florida, Gainesville, FL, USA · Department of Radiology, University of Florida, Jacksonville, FL, USA · Research Computing, University of Florida, Gainesville, FL, USA · Department of Radiology, UC San Francisco, San Francisco, CA, USA · Department of Medicine, University of Florida, Gainesville, FL, USA
Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly when aiming to unify these capabilities within a single, generalizable framework. To address this challenge, we proposed MedVL-SAM2, a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multi-paradigm segmentation, including semantic, referring, and interactive segmentation. MedVL-SAM2 integrates image-level reasoning and pixel-level perception through a cohesive architecture tailored for 3D medical imaging, and incorporates a SAM2-based volumetric segmentation module to enable precise multi-granular spatial reasoning. The model is trained in a multi-stage pipeline: it is first pre-trained on a large-scale corpus of 3D CT image-text pairs to align volumetric visual features with radiology-language embeddings. It is then jointly optimized with both language-understanding and segmentation objectives using a comprehensive 3D CT segmentation dataset. This joint training enables flexible interaction via language, point, or box prompts, thereby unifying high-level visual reasoning with spatially precise localization. Our unified architecture delivers state-of-the-art performance across report generation, VQA, and multiple 3D segmentation tasks. Extensive analyses further show that the model provides reliable 3D visual grounding, controllable interactive segmentation, and robust cross-modal reasoning, demonstrating that high-level semantic reasoning and precise 3D localization can be jointly achieved within a unified 3D medical VLM.
Figures & tables
Fig. 1: Comparison of the proposed model to traditional text-response-only medical vision-language models (VLMs) and mask-response-only medical vision foundation models (VFMs).
Method
RG
VQA
SS
RS
IS
BiomedParse [ 7 ]
×
×
✓
✓
×
medSAM2 [ 8 ]
×
×
✓
×
✓
CT-CHAT [ 4 ]
✓
✓
×
×
×
Med3DVLM [ 3 ]
✓
✓
×
×
×
Med-2E3 [ 9 ]
✓
✓
×
×
×
MS-VLM [ 10 ]
✓
✓
×
×
×
TABLE I: Capability comparison between our model and current state-of-the-art methods across report generation (RG), visual question answering (VQA), semantic segmentation (SS), referring segmentation (RS), and interactive segmentation (IS).
Fig. 2: Overview of the proposed architecture, which integrates a LLaVA-style VLM with a SAM2-based segmentation module. The VLM processes 3D volumes and generates text autoregressively. When a [SEG] token is produced, its hidden state is extracted and fed into SAM2’s prompt encoder, where it is fused with optional visual prompts (points or boxes) to generate the final segmentation mask.
Module
Parameters
Size
Vision Encoder
198.05M
377.76 MB
Projection Layer
7.09M
13.52 MB
LLM Backbone
3.39B
6.33 GB
Segmentation Module
224.43M
428.07 MB
Total
3.82B
7.15 GB
TABLE II: Module parameters and model size
Fig. 3: Qualitative comparison of the proposed method with CT-CHAT on the report generation task. Coronal, axial, and sagittal views of the corresponding CT scans are shown. Different aspects of the reports are color-coded for improved readability.
Fig. 4: Comparison between the proposed methods to CT-CHAT on VQA tasks, including Long Answer, Short Answer, and Multiple choice subsets. Critical clinical information was highlighted in red. For CT-CHAT, the input question additionally contained a special token ( [short_answer] , [long_answer] , [multiple_choice] ) to indicate the specific subset to evaluate.
Method
BLEU-1
ROUGE
METEOR
CIDEr
BERTScore
LLaVA [ 13 ]
18.47
10.34
11.22
0.050
30.55
CXR-LLaVA [ 47 ]
9.12
5.05
2.86
0.049
39.18
LLaVA-Med [ 27 ]
9.01
5.24
2.24
0.006
40.65
M3D [ 5 ]
14.49
19.25
14.11
0.081
84.12
CT-CHAT [ 4 ]
43.64
31.57
23.06
0.221
88.12
Ours
41.85
34.59
22.15
0.237
89.33
TABLE III: Performance comparison on report generation.
Fig. 5: Qualitative comparison of the proposed method with M3D on referring and semantic segmentation. Interactive segmentation results using bounding-box prompts are also included for comparison (ACT-1K dataset). The liver is shown in red, the kidney in cyan, the spleen in orange, and the pancreas in blue.
Fig. 6: Comparison of the proposed method with M3D on referring and semantic segmentation. Interactive segmentation results using bounding-box prompts are also included for comparison (CTOrg dataset shown). Liver is shown in magenta, bladder in green, lung in brown, kidney in purple, bone in red, and brain in blue.
Task
Long Answer
Short Answer
Multiple Choice
BLEU
ROUGE
METEOR
CIDEr
BERTScore
BLEU
ROUGE
METEOR
CIDEr
BERTScore
Accuracy
CT-CHAT [ 4 ]
47.020
48.450
28.200
2.910
90.140
27.460
45.590
15.470
1.652
90.120
83.730
Ours
45.941
48.983
28.218
2.861
92.760
24.076
57.428
14.026
1.771
92.503
89.741
TABLE IV: Comparison between CT-CHAT and our model on long-answer, short-answer, and multiple-choice VQA tasks.
Method
CTOrg
ACT-1K
TotalSegmentator
SegVol [ 40 ]
77.78
79.06
44.28
M3D [ 5 ] (semantic)
81.27
73.64
58.30
M3D [ 5 ] (referring)
83.49
73.63
58.50
Ours (semantic)
88.04
88.27
66.23
Ours (referring)
87.93
88.34
66.50
TABLE V: Results on the referring segmentation task.
Method
CTOrg
ACT-1K
TotalSegmentator
Ours (semantic)
88.04
88.27
66.23
Ours (referring)
87.93
88.34
66.50
Ours (bbox prompt)
88.76
88.71
70.51
Ours (points prompt)
88.52
88.49
69.88
TABLE VI: Results on the interactive segmentation task.
Vision Encoder
Projection Layer
CTOrg
ACT-1K
Vanilla CLIP [ 43 ]
Linear
30.18
44.65
M3D-CLIP [ 5 ]
Linear
61.99
74.58
M3D-CLIP [ 5 ]
MLP
77.78
79.83
M3D-CLIP [ 5 ]
MLP-Mixer
88.04
88.27
TABLE VII: Ablation studies on different choices of vision encoders and projection layers.
Training strategy
BLEU-1
ROUGE
METEOR
CIDEr
One-stage
38.58
32.53
20.89
0.2095
Multi-stage
41.85
34.59
22.15
0.2371
TABLE VIII: Ablation studies on different choices of training strategies.
Fig. 7: Prompt templates used for report generation tasks.
Fig. 8: Prompt templates used for referring segmentation tasks.
Fig. 9: Prompt templates used for semantic segmentation tasks.
Fig. 10: Anatomical descriptions used as free-form text prompts for referring segmentation tasks. These descriptions replace the {} placeholder in the templates shown in Figure 8 .