MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
Authors: Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
Organizations: Department of Biomedical Engineering, University of Florida, Gainesville, FL, USA · Department of Radiology, University of Florida, Jacksonville, FL, USA · Research Computing, University of Florida, Gainesville, FL, USA · Department of Radiology, UC San Francisco, San Francisco, CA, USA · Department of Medicine, University of Florida, Gainesville, FL, USA
Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly when aiming to unify these capabilities within a single, generalizable framework. To address this challenge, we proposed MedVL-SAM2, a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multi-paradigm segmentation, including semantic, referring, and interactive segmentation. MedVL-SAM2 integrates image-level reasoning and pixel-level perception through a cohesive architecture tailored for 3D medical imaging, and incorporates a SAM2-based volumetric segmentation module to enable precise multi-granular spatial reasoning. The model is trained in a multi-stage pipeline: it is first pre-trained on a large-scale corpus of 3D CT image-text pairs to align volumetric visual features with radiology-language embeddings. It is then jointly optimized with both language-understanding and segmentation objectives using a comprehensive 3D CT segmentation dataset. This joint training enables flexible interaction via language, point, or box prompts, thereby unifying high-level visual reasoning with spatially precise localization. Our unified architecture delivers state-of-the-art performance across report generation, VQA, and multiple 3D segmentation tasks. Extensive analyses further show that the model provides reliable 3D visual grounding, controllable interactive segmentation, and robust cross-modal reasoning, demonstrating that high-level semantic reasoning and precise 3D localization can be jointly achieved within a unified 3D medical VLM.
Figures & tables
Fig. 1: Comparison of the proposed model to traditional text-response-only medical vision-language models (VLMs) and mask-response-only medical vision foundation models (VFMs).
Method
RG
VQA
SS
RS
IS
BiomedParse [ 7 ]
×
×
✓
✓
×
medSAM2 [ 8 ]
×
×
✓
×
✓
CT-CHAT [ 4 ]
✓
✓
×
×
×
Med3DVLM [ 3 ]
✓
✓
×
×
×
Med-2E3 [ 9 ]
✓
✓
×
×
×
MS-VLM [ 10 ]
✓
✓
×
×
×
TABLE I: Capability comparison between our model and current state-of-the-art methods across report generation (RG), visual question answering (VQA), semantic segmentation (SS), referring segmentation (RS), and interactive segmentation (IS).
Fig. 2: Overview of the proposed architecture, which integrates a LLaVA-style VLM with a SAM2-based segmentation module. The VLM processes 3D volumes and generates text autoregressively. When a [SEG] token is produced, its hidden state is extracted and fed into SAM2’s prompt encoder, where it is fused with optional visual prompts (points or boxes) to generate the final segmentation mask.
Module
Parameters
Size
Vision Encoder
198.05M
377.76 MB
Projection Layer
7.09M
13.52 MB
LLM Backbone
3.39B
6.33 GB
Segmentation Module
224.43M
428.07 MB
Total
3.82B
7.15 GB
TABLE II: Module parameters and model size
Fig. 3: Qualitative comparison of the proposed method with CT-CHAT on the report generation task. Coronal, axial, and sagittal views of the corresponding CT scans are shown. Different aspects of the reports are color-coded for improved readability.
Fig. 4: Comparison between the proposed methods to CT-CHAT on VQA tasks, including Long Answer, Short Answer, and Multiple choice subsets. Critical clinical information was highlighted in red. For CT-CHAT, the input question additionally contained a special token ( [short_answer] , [long_answer] , [multiple_choice] ) to indicate the specific subset to evaluate.
Method
BLEU-1
ROUGE
METEOR
CIDEr
BERTScore
LLaVA [ 13 ]
18.47
10.34
11.22
0.050
30.55
CXR-LLaVA [ 47 ]
9.12
5.05
2.86
0.049
39.18
LLaVA-Med [ 27 ]
9.01
5.24
2.24
0.006
40.65
M3D [ 5 ]
14.49
19.25
14.11
0.081
84.12
CT-CHAT [ 4 ]
43.64
31.57
23.06
0.221
88.12
Ours
41.85
34.59
22.15
0.237
89.33
TABLE III: Performance comparison on report generation.
Fig. 5: Qualitative comparison of the proposed method with M3D on referring and semantic segmentation. Interactive segmentation results using bounding-box prompts are also included for comparison (ACT-1K dataset). The liver is shown in red, the kidney in cyan, the spleen in orange, and the pancreas in blue.
Fig. 6: Comparison of the proposed method with M3D on referring and semantic segmentation. Interactive segmentation results using bounding-box prompts are also included for comparison (CTOrg dataset shown). Liver is shown in magenta, bladder in green, lung in brown, kidney in purple, bone in red, and brain in blue.
Task
Long Answer
Short Answer
Multiple Choice
BLEU
ROUGE
METEOR
CIDEr
BERTScore
BLEU
ROUGE
METEOR
CIDEr
BERTScore
Accuracy
CT-CHAT [ 4 ]
47.020
48.450
28.200
2.910
90.140
27.460
45.590
15.470
1.652
90.120
83.730
Ours
45.941
48.983
28.218
2.861
92.760
24.076
57.428
14.026
1.771
92.503
89.741
TABLE IV: Comparison between CT-CHAT and our model on long-answer, short-answer, and multiple-choice VQA tasks.
Method
CTOrg
ACT-1K
TotalSegmentator
SegVol [ 40 ]
77.78
79.06
44.28
M3D [ 5 ] (semantic)
81.27
73.64
58.30
M3D [ 5 ] (referring)
83.49
73.63
58.50
Ours (semantic)
88.04
88.27
66.23
Ours (referring)
87.93
88.34
66.50
TABLE V: Results on the referring segmentation task.
Method
CTOrg
ACT-1K
TotalSegmentator
Ours (semantic)
88.04
88.27
66.23
Ours (referring)
87.93
88.34
66.50
Ours (bbox prompt)
88.76
88.71
70.51
Ours (points prompt)
88.52
88.49
69.88
TABLE VI: Results on the interactive segmentation task.
Vision Encoder
Projection Layer
CTOrg
ACT-1K
Vanilla CLIP [ 43 ]
Linear
30.18
44.65
M3D-CLIP [ 5 ]
Linear
61.99
74.58
M3D-CLIP [ 5 ]
MLP
77.78
79.83
M3D-CLIP [ 5 ]
MLP-Mixer
88.04
88.27
TABLE VII: Ablation studies on different choices of vision encoders and projection layers.
Training strategy
BLEU-1
ROUGE
METEOR
CIDEr
One-stage
38.58
32.53
20.89
0.2095
Multi-stage
41.85
34.59
22.15
0.2371
TABLE VIII: Ablation studies on different choices of training strategies.
Fig. 7: Prompt templates used for report generation tasks.
Fig. 8: Prompt templates used for referring segmentation tasks.
Fig. 9: Prompt templates used for semantic segmentation tasks.
Fig. 10: Anatomical descriptions used as free-form text prompts for referring segmentation tasks. These descriptions replace the {} placeholder in the templates shown in Figure 8 .
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric imaging remains hindered by prohibitive annotation costs and dataset opacity. Current data formats, predominantly consisting of rigid Visual Question Answering (VQA) pairs or unstructured final clinical reports, typically fail to capture explicit clinical reasoning. To address this limitation, we introduce a large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm. Inspired by the genuine diagnostic workflow of radiologists, this paradigm models visual cognition by decomposing the complex 3D reading process, translating global clinical priors into fine-grained, per-slice observations that are subsequently synthesized into an interpretable Chain-of-Thought (CoT). Crucially, this synthesized reasoning framework enforces essential clinical principles: sequential spatial tracking, multi-slice spatial awareness for artifact mitigation, and differential exclusion. To validate this approach, we instruction-tune a standard 2D-pretrained MLLM baseline using the synthesized data to enhance its volumetric comprehension. Comprehensive evaluations across multiple 3D medical benchmarks demonstrate that our method yields significant performance improvements over the 2D baseline. Furthermore, the resulting model exhibits robust spatial reasoning capabilities and rivals resource-intensive native 3D architectures, effectively bridging the performance gap. Ultimately, this data-centric strategy unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training. The complete repository, including datasets and training workflows, is publicly available at https://github.com/2020420145009/hounsfield.
Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5
University of International Relations, Beijing, China · University of Trento, Trento, Italy
Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special <SEG> token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the <SEG> hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable <SEG> prompt states, monitored by the Davies--Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.
Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizability and demand extensive, time-consuming manual annotation for new clinical application. Here, we propose MedSAM-3, a text promptable medical segmentation model for medical image and video segmentation. By fine-tuning the Segment Anything Model (SAM) 3 architecture on medical images paired with semantic conceptual labels, our MedSAM-3 enables medical Promptable Concept Segmentation (PCS), allowing precise targeting of anatomical structures via open-vocabulary text descriptions rather than solely geometric prompts. We further introduce the MedSAM-3 Agent, a framework that integrates Multimodal Large Language Models (MLLMs) to perform complex reasoning and iterative refinement in an agent-in-the-loop workflow. Comprehensive experiments across diverse medical imaging modalities, including X-ray, MRI, Ultrasound, CT, and video, demonstrate that our approach significantly outperforms existing specialist and foundation models. We will release our code and model at https://github.com/Joey-S-Liu/MedSAM3.
Anglin Liu, Xu R. Cao, Yifan Shen +4
The Hong Kong University of Science and Technology (Guangzhou) · University of Illinois Urbana-Champaign · Southeast University +1