Organizations: School of Computer Science and Information Engineering, Hefei University of Technology · Nanjing University · Guangdong University of Finance and Economics · East China Normal University · Tianxi AI Technology Platform, Lenovo
While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, an efficient medical multimodal multi-agent collaboration framework that dynamically prunes both nodes and edges from the communication topology to enhance reasoning ability and token efficiency. Specifically, we first formulate the diagnostic process as a heterogeneous communication graph, where nodes represent specialist agents from various departments and edges capture intra- and inter-departmental interactions. Building on this graph, we introduce two sparsification mechanisms to enable adaptive collaborative evolution: (1) Heterogeneous Node Sparsification, which eliminates task-irrelevant specialist agents irrelevant to the current multimodal question via reinforcement learning-driven topological optimization, and (2) Heterogeneous Edge Sparsification, which selectively retains only the most diagnostically salient intra- and inter-departmental connections by jointly optimizing task performance and topological complexity. Extensive medical VQA experiments under full-set and few-shot training settings prove MedPrune surpasses multi-agent baselines and boosts token efficiency with strong adversarial robustness.
Figures & tables
Figure 1: Comparison of our approach with existing multi-agent methods. Fixed-Template Evolving relies on static and predefined communication templates, which leads to degraded diagnostic performance due to noisy redundant discussions. Adaptive Evolving can also yield a smaller communication subgraph in the final stage, yet its iterative topology refinement still incur substantial token overhead, failing to achieve true medical efficiency gains. In contrast, our MedPrune dynamically prunes irrelevant medical agents and redundant edges simultaneously, enabling token-efficient multi-department diagnostic reasoning (Rounds T2≪T1 ).
Figure 2: Overview of MedPrune . (1) Heterogeneous Node Sparsification removes specialist agents irrelevant to the medical question. (2) Heterogeneous Edge Sparsification further prunes redundant intra- and inter-departmental connections via jointly maximizing task performance and penalizes topological complexity. Due to the space limitation, we only draw four spatial and temporal edge connections for one medical department.
Training Paradigm
Method
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Full-set Training
Zero-shot
SP
7.65
18.09
62.58
29.64
29.49
84.08
72.36
61.98±0.7
CoT
15.25
28.57
69.13
24.60
34.39
86.02
71.07
63.83±1.1
Multimodal RAG
MMed-RAG
15.45
26.01
69.36
33.94
36.19
86.21
82.31
68.24±0.5
Table 1: The performance of full-set and few-shot training with Qwen3-VL (8B) . Due to space limitation, the results across different backbones ( HuatuoGPT-Vision (7B) , Qwen3.6-Plus (397B) ) and few-shot settings with 20/40 training samples are provided in Appendix D.1 . The t-tests demonstrate the improvements are statistically significant with p < 0.05 level. Best results are in bold and second-best results are underlined.
Dataset →
MIMIC-CXR
IU-Xray
Average
Method ↓
Qwen3-VL (8B)
MedPrune
42.26
91.07
66.67
w/o NS
38.59 {{\color[rgb]{1,0.5,0}(\downarrow 3.67)}}
86.41 {{\color[rgb]{1,0.5,0}(\downarrow 4.66)}}
62.50 {{\color[rgb]{1,0.5,0}(\downarrow 4.17)}}
w/o ES
38.55 {{\color[rgb]{1,0.5,0}(\downarrow 3.71)}}
85.44 {{\color[rgb]{1,0.5,0}(\downarrow 5.63)}}
62.00 {{\color[rgb]{1,0.5,0}(\downarrow 4.67)}}
w/o norm
41.24 {{\color[rgb]{1,0.5,0}(\downarrow 1.02)}}
86.99 {{\color[rgb]{1,0.5,0}(\downarrow 4.08)}}
64.12 {{\color[rgb]{1,0.5,0}(\downarrow 2.55)}}
Table 2: Ablation study of MedPrune . "NS", "ES", "PP" and "DS" are abbreviations for node sparsification, edge sparsification, progressive edge pruning and department selection, respectively.
Figure 3: Token efficiency ( F /#Tokens) comparison across various medical multi-agent frameworks.
Figure 4: Visualization of intra - (Left) and inter - (Right) department communication edge weight evolution across multi - agent interaction rounds (Best viewed in color).
Figure 5: Model performance retention under adversarial attacks (prompt and response perturbations).
Dataset
Training Settings
Training Size
Testing Size
MIMIC-CXR
Full-set
8.3k
2.0k
Few-shot
10/20/40
10.3k
IU-Xray
Full-set
2.1k
0.5k
Few-shot
10/20/40
2.6k
OmniMedVQA
Full-set
8.0k
2.0k
Few-shot
10/20/40
10.0k
Table 3: Statistics of three medical multimodal datasets under full-set and few-shot settings.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
G
heterogeneous multi-agent diagnostic graph
G^
Communication graph after Sparsification
V
set of all specialist agent nodes
ET
set of temporal edges
ES
set of spatial edges
M
memory states of all agents
Appendix
Table 4: Notations and Descriptions in MedPrune .
Figure 6: The influence of different agent configurations using Qwen3-VL (8B) and HuatuoGPT-Vision (7B).
Figure 7: The influence of different node and edge pruning rates using Qwen3-VL (8B) and HuatuoGPT-Vision (7B).
Dataset →
MIMIC-CXR
IU-Xray
Average
Settings ↓
Base model: Qwen3-VL (8B)
T=2,δ=0.1
42.26
91.07
66.67
T=4,δ=0.1
41.48
89.51
65.49
T=6,δ=0.1
41.64
90.29
65.97
T=8,δ=0.1
41.72
90.10
65.91
Appendix
Table 5: Hyperparameter experiments for MedPrune under different numbers of communication rounds ( T ) and noise levels ( δ ).
Dataset →
MIMIC-CXR
IU-Xray
Average
Settings ↓
Base model: Qwen3-VL (8B)
#20
43.08
89.71
66.40
#40
42.26
91.07
66.67
#60
41.74
89.51
65.63
#80
41.60
90.29
65.94
Appendix
Table 6: Results under various training settings using Qwen3-VL (8B) and HuatuoGPT-Vision (7B). “#40” indicates that the number of training samples is 40.
Figure 8: Prompt templates for all specialist agents, including chief physicians and regular doctors across 11 clinical departments. Considering the formatting, we have placed other templates in Fig. 9 .
Figure 9: Prompt templates for other specialist agents.
Training Paradigm
Method
Sample Scale
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Few-shot Training (20/40 samples)
Zero-shot
SP
20
7.62
17.92
62.58
29.37
29.37
84.49
69.87
61.24 ±0.7
40
7.65
17.95
62.61
29.52
29.43
84.88
70.13
61.48 ±1.0
CoT
20
13.52
31.15
68.47
23.18
34.08
85.74
70.70
63.51 ±0.8
Appendix
Table 7: The performance of few-shot settings with 20/40 training samples with Qwen3-VL (8B) . The t-tests demonstrate the improvements are statistically significant with p < 0.05 level.
Training Paradigm
Method
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Full-set Training
Zero-shot
SP
11.75
24.98
64.43
21.27
30.61
74.17
70.77
58.52 ±0.7
CoT
18.76
33.97
70.36
28.67
37.94
76.12
72.36
62.14 ±0.9
Multimodal RAG
MMed-RAG
20.59
30.82
71.16
30.60
38.29
84.08
84.99
69.12 ±0.8
Appendix
Table 8: Full-set training performance with HuatuoGPT-Vision (7B) . The t-tests demonstrate the improvements are statistically significant with p<0.05 level. Best results are in bold and second-best results are underlined.
Training Paradigm
Method
Sample Scale
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Few-shot Training
Zero-shot
SP
10
11.92
24.88
64.15
21.26
30.55
75.73
70.55
58.94 ±0.6
20
11.03
24.77
64.53
21.25
30.39
74.50
70.48
58.46 ±0.8
40
11.90
24.76
64.13
21.22
30.50
74.65
70.57
58.57 ±0.5
Appendix
Table 9: Few-shot training performance with HuatuoGPT-Vision (7B) . The t-tests demonstrate the improvements are statistically significant with p<0.05 level. Best results are in bold and second-best results are underlined.
Training Paradigm
Method
Sample Size
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Full-set Training
Zero-shot
SP
/
18.70
30.38
69.46
28.52
36.77
88.54
87.08
70.80 ±0.7
CoT
/
19.14
30.62
70.76
29.43
37.49
87.38
89.86
71.58 ±1.0
Multimodal RAG
MMed-RAG
/
21.18
32.35
71.04
31.22
38.95
90.10
93.74
74.26 ±0.8
Appendix
Table 10: The performance of full-set and few-shot training is reported using the closed - source Qwen3.6-Plus (397B) accessed via its API. Since the API does not support gradient backpropagation or parameter updates, model fine-tuning is infeasible. We therefore restrict our comparison to training-free baselines. The t-tests demonstrate the improvements are statistically significant with p<0.05 level. Bold black indicates the best results, while underlined is the second ranked results.
Dataset →
MIMIC-CXR
IU-Xray
OmniMedVQA
Models ↓
#Tokens
F
#Tokens
F
#Tokens
F
CARE
11.11k
40.08
6.46k
89.90
12.79k
92.80
AdaCoMed
10.24k
41.46
8.04k
90.10
13.19k
92.45
MedAgent-Pro
8.59k
41.42
6.31k
88.35
13.03k
91.01
MMedAgent-RL
11.24k
41.49
6.09k
88.93
14.12k
91.80
M 3 Prune
8.23k
40.09
6.02k
89.13
12.92k
90.51
Appendix
Table 11: Token consumption of per multi-modal medical question and dataset metrics ( F ) on three datasets.
Figure 10: Prompt template for input prompt attack.
Figure 11: Prompt template for response prompt attack.
Method
Tune Backbone
Total GPU Hours (H)
Peak VRAM (GB)
Train Convergence Epochs
Train Tokens
CARE / AdaCoMed MedAgent-Pro
✗
0
< 10
0
0
MMedAgent-RL / M 3 Prune ARG-Designer
Topology Module
17 / 5 / 7
49 / 32 / 29
8 / 2 / 5
18.6k / 14.3k / 16.5k
MedPrune (Ours)
Topology Module
1.5
18
2
10.2k
Appendix
Table 12: Training average overhead of all multi-agent baselines on single GPU on three datasets. “Tune Backbone”: ✗: Core vision-language transformer backbone fully frozen with no trainable extra modules; Topology Module: Only multi-agent communication graph parameters optimized, vision-language backbone frozen. The training process of multi-agent baselines is one-time , but the model used in the inference process is reused multiple times .
Prompt Variant
MIMIC-CXR
IU-Xray
OmniMedVQA
Average
Prompt-Ours (Main)
42.26
91.07
96.36
76.56
Prompt-Detailed
42.01
90.87
95.91
76.26
Prompt-Simplified
41.75
90.68
95.46
75.96
Prompt-Reversed
41.32
90.49
94.81
75.54
Appendix
Table 13: Prompt sensitivity of MedPrune . The well-trained topological pruning weights are fixed, only the agent prompt templates are replaced during inference.
Dataset →
MIMIC-CXR → IU-Xray
PMC - OA → MIMIC-CXR
Average
Method ↓
AdaCoMed
85.05
36.95
61.00
MMedAgent-RL
82.91
35.98
59.45
M 3 Prune
85.44
36.99
61.22
ARG-Designer †
84.46
37.02
60.74
Ours (MedPrune)
87.96
38.96
63.46
Appendix
Table 14: Cross-Dataset Transfer results under full-set training based on Qwen3 - VL (8B) backbone. MIMIC-CXR→IU-Xray reports accuracy (%); PMC-OA→MIMIC-CXR reports the average of BLEU, ROUGE-L, BERT-F1 and METEOR (Avg.).
Figure 12: Case study of Round 1 Discussion.
Figure 13: Case study of Round 2 Discussion and Final Decision.