Organizations: School of Computer Science and Information Engineering, Hefei University of Technology · Nanjing University · Guangdong University of Finance and Economics · East China Normal University · Tianxi AI Technology Platform, Lenovo
While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, an efficient medical multimodal multi-agent collaboration framework that dynamically prunes both nodes and edges from the communication topology to enhance reasoning ability and token efficiency. Specifically, we first formulate the diagnostic process as a heterogeneous communication graph, where nodes represent specialist agents from various departments and edges capture intra- and inter-departmental interactions. Building on this graph, we introduce two sparsification mechanisms to enable adaptive collaborative evolution: (1) Heterogeneous Node Sparsification, which eliminates task-irrelevant specialist agents irrelevant to the current multimodal question via reinforcement learning-driven topological optimization, and (2) Heterogeneous Edge Sparsification, which selectively retains only the most diagnostically salient intra- and inter-departmental connections by jointly optimizing task performance and topological complexity. Extensive medical VQA experiments under full-set and few-shot training settings prove MedPrune surpasses multi-agent baselines and boosts token efficiency with strong adversarial robustness.
Figures & tables
Figure 1: Comparison of our approach with existing multi-agent methods. Fixed-Template Evolving relies on static and predefined communication templates, which leads to degraded diagnostic performance due to noisy redundant discussions. Adaptive Evolving can also yield a smaller communication subgraph in the final stage, yet its iterative topology refinement still incur substantial token overhead, failing to achieve true medical efficiency gains. In contrast, our MedPrune dynamically prunes irrelevant medical agents and redundant edges simultaneously, enabling token-efficient multi-department diagnostic reasoning (Rounds T2≪T1 ).
Figure 2: Overview of MedPrune . (1) Heterogeneous Node Sparsification removes specialist agents irrelevant to the medical question. (2) Heterogeneous Edge Sparsification further prunes redundant intra- and inter-departmental connections via jointly maximizing task performance and penalizes topological complexity. Due to the space limitation, we only draw four spatial and temporal edge connections for one medical department.
Training Paradigm
Method
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Full-set Training
Zero-shot
SP
7.65
18.09
62.58
29.64
29.49
84.08
72.36
61.98±0.7
CoT
15.25
28.57
69.13
24.60
34.39
86.02
71.07
63.83±1.1
Multimodal RAG
MMed-RAG
15.45
26.01
69.36
33.94
36.19
86.21
82.31
68.24±0.5
Table 1: The performance of full-set and few-shot training with Qwen3-VL (8B) . Due to space limitation, the results across different backbones ( HuatuoGPT-Vision (7B) , Qwen3.6-Plus (397B) ) and few-shot settings with 20/40 training samples are provided in Appendix D.1 . The t-tests demonstrate the improvements are statistically significant with p < 0.05 level. Best results are in bold and second-best results are underlined.
Dataset →
MIMIC-CXR
IU-Xray
Average
Method ↓
Qwen3-VL (8B)
MedPrune
42.26
91.07
66.67
w/o NS
38.59 {{\color[rgb]{1,0.5,0}(\downarrow 3.67)}}
86.41 {{\color[rgb]{1,0.5,0}(\downarrow 4.66)}}
62.50 {{\color[rgb]{1,0.5,0}(\downarrow 4.17)}}
w/o ES
38.55 {{\color[rgb]{1,0.5,0}(\downarrow 3.71)}}
85.44 {{\color[rgb]{1,0.5,0}(\downarrow 5.63)}}
62.00 {{\color[rgb]{1,0.5,0}(\downarrow 4.67)}}
w/o norm
41.24 {{\color[rgb]{1,0.5,0}(\downarrow 1.02)}}
86.99 {{\color[rgb]{1,0.5,0}(\downarrow 4.08)}}
64.12 {{\color[rgb]{1,0.5,0}(\downarrow 2.55)}}
Table 2: Ablation study of MedPrune . "NS", "ES", "PP" and "DS" are abbreviations for node sparsification, edge sparsification, progressive edge pruning and department selection, respectively.
Figure 3: Token efficiency ( F /#Tokens) comparison across various medical multi-agent frameworks.
Figure 4: Visualization of intra - (Left) and inter - (Right) department communication edge weight evolution across multi - agent interaction rounds (Best viewed in color).
Figure 5: Model performance retention under adversarial attacks (prompt and response perturbations).
Dataset
Training Settings
Training Size
Testing Size
MIMIC-CXR
Full-set
8.3k
2.0k
Few-shot
10/20/40
10.3k
IU-Xray
Full-set
2.1k
0.5k
Few-shot
10/20/40
2.6k
OmniMedVQA
Full-set
8.0k
2.0k
Few-shot
10/20/40
10.0k
Table 3: Statistics of three medical multimodal datasets under full-set and few-shot settings.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
G
heterogeneous multi-agent diagnostic graph
G^
Communication graph after Sparsification
V
set of all specialist agent nodes
ET
set of temporal edges
ES
set of spatial edges
M
memory states of all agents
Appendix
Table 4: Notations and Descriptions in MedPrune .
Figure 6: The influence of different agent configurations using Qwen3-VL (8B) and HuatuoGPT-Vision (7B).
Figure 7: The influence of different node and edge pruning rates using Qwen3-VL (8B) and HuatuoGPT-Vision (7B).
Dataset →
MIMIC-CXR
IU-Xray
Average
Settings ↓
Base model: Qwen3-VL (8B)
T=2,δ=0.1
42.26
91.07
66.67
T=4,δ=0.1
41.48
89.51
65.49
T=6,δ=0.1
41.64
90.29
65.97
T=8,δ=0.1
41.72
90.10
65.91
Appendix
Table 5: Hyperparameter experiments for MedPrune under different numbers of communication rounds ( T ) and noise levels ( δ ).
Dataset →
MIMIC-CXR
IU-Xray
Average
Settings ↓
Base model: Qwen3-VL (8B)
#20
43.08
89.71
66.40
#40
42.26
91.07
66.67
#60
41.74
89.51
65.63
#80
41.60
90.29
65.94
Appendix
Table 6: Results under various training settings using Qwen3-VL (8B) and HuatuoGPT-Vision (7B). “#40” indicates that the number of training samples is 40.
Figure 8: Prompt templates for all specialist agents, including chief physicians and regular doctors across 11 clinical departments. Considering the formatting, we have placed other templates in Fig. 9 .
Figure 9: Prompt templates for other specialist agents.
Training Paradigm
Method
Sample Scale
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Few-shot Training (20/40 samples)
Zero-shot
SP
20
7.62
17.92
62.58
29.37
29.37
84.49
69.87
61.24 ±0.7
40
7.65
17.95
62.61
29.52
29.43
84.88
70.13
61.48 ±1.0
CoT
20
13.52
31.15
68.47
23.18
34.08
85.74
70.70
63.51 ±0.8
Appendix
Table 7: The performance of few-shot settings with 20/40 training samples with Qwen3-VL (8B) . The t-tests demonstrate the improvements are statistically significant with p < 0.05 level.
Training Paradigm
Method
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Full-set Training
Zero-shot
SP
11.75
24.98
64.43
21.27
30.61
74.17
70.77
58.52 ±0.7
CoT
18.76
33.97
70.36
28.67
37.94
76.12
72.36
62.14 ±0.9
Multimodal RAG
MMed-RAG
20.59
30.82
71.16
30.60
38.29
84.08
84.99
69.12 ±0.8
Appendix
Table 8: Full-set training performance with HuatuoGPT-Vision (7B) . The t-tests demonstrate the improvements are statistically significant with p<0.05 level. Best results are in bold and second-best results are underlined.
Training Paradigm
Method
Sample Scale
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Few-shot Training
Zero-shot
SP
10
11.92
24.88
64.15
21.26
30.55
75.73
70.55
58.94 ±0.6
20
11.03
24.77
64.53
21.25
30.39
74.50
70.48
58.46 ±0.8
40
11.90
24.76
64.13
21.22
30.50
74.65
70.57
58.57 ±0.5
Appendix
Table 9: Few-shot training performance with HuatuoGPT-Vision (7B) . The t-tests demonstrate the improvements are statistically significant with p<0.05 level. Best results are in bold and second-best results are underlined.
Training Paradigm
Method
Sample Size
MIMIC-CXR
IU-Xray
Omni- MedVQA
Average
BLEU
ROUGE-L
BERT-F1
METEOR
Avg.
Acc
Acc
Full-set Training
Zero-shot
SP
/
18.70
30.38
69.46
28.52
36.77
88.54
87.08
70.80 ±0.7
CoT
/
19.14
30.62
70.76
29.43
37.49
87.38
89.86
71.58 ±1.0
Multimodal RAG
MMed-RAG
/
21.18
32.35
71.04
31.22
38.95
90.10
93.74
74.26 ±0.8
Appendix
Table 10: The performance of full-set and few-shot training is reported using the closed - source Qwen3.6-Plus (397B) accessed via its API. Since the API does not support gradient backpropagation or parameter updates, model fine-tuning is infeasible. We therefore restrict our comparison to training-free baselines. The t-tests demonstrate the improvements are statistically significant with p<0.05 level. Bold black indicates the best results, while underlined is the second ranked results.
Dataset →
MIMIC-CXR
IU-Xray
OmniMedVQA
Models ↓
#Tokens
F
#Tokens
F
#Tokens
F
CARE
11.11k
40.08
6.46k
89.90
12.79k
92.80
AdaCoMed
10.24k
41.46
8.04k
90.10
13.19k
92.45
MedAgent-Pro
8.59k
41.42
6.31k
88.35
13.03k
91.01
MMedAgent-RL
11.24k
41.49
6.09k
88.93
14.12k
91.80
M 3 Prune
8.23k
40.09
6.02k
89.13
12.92k
90.51
Appendix
Table 11: Token consumption of per multi-modal medical question and dataset metrics ( F ) on three datasets.
Figure 10: Prompt template for input prompt attack.
Figure 11: Prompt template for response prompt attack.
Method
Tune Backbone
Total GPU Hours (H)
Peak VRAM (GB)
Train Convergence Epochs
Train Tokens
CARE / AdaCoMed MedAgent-Pro
✗
0
< 10
0
0
MMedAgent-RL / M 3 Prune ARG-Designer
Topology Module
17 / 5 / 7
49 / 32 / 29
8 / 2 / 5
18.6k / 14.3k / 16.5k
MedPrune (Ours)
Topology Module
1.5
18
2
10.2k
Appendix
Table 12: Training average overhead of all multi-agent baselines on single GPU on three datasets. “Tune Backbone”: ✗: Core vision-language transformer backbone fully frozen with no trainable extra modules; Topology Module: Only multi-agent communication graph parameters optimized, vision-language backbone frozen. The training process of multi-agent baselines is one-time , but the model used in the inference process is reused multiple times .
Prompt Variant
MIMIC-CXR
IU-Xray
OmniMedVQA
Average
Prompt-Ours (Main)
42.26
91.07
96.36
76.56
Prompt-Detailed
42.01
90.87
95.91
76.26
Prompt-Simplified
41.75
90.68
95.46
75.96
Prompt-Reversed
41.32
90.49
94.81
75.54
Appendix
Table 13: Prompt sensitivity of MedPrune . The well-trained topological pruning weights are fixed, only the agent prompt templates are replaced during inference.
Dataset →
MIMIC-CXR → IU-Xray
PMC - OA → MIMIC-CXR
Average
Method ↓
AdaCoMed
85.05
36.95
61.00
MMedAgent-RL
82.91
35.98
59.45
M 3 Prune
85.44
36.99
61.22
ARG-Designer †
84.46
37.02
60.74
Ours (MedPrune)
87.96
38.96
63.46
Appendix
Table 14: Cross-Dataset Transfer results under full-set training based on Qwen3 - VL (8B) backbone. MIMIC-CXR→IU-Xray reports accuracy (%); PMC-OA→MIMIC-CXR reports the average of BLEU, ROUGE-L, BERT-F1 and METEOR (Avg.).
Figure 12: Case study of Round 1 Discussion.
Figure 13: Case study of Round 2 Discussion and Final Decision.
Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle with medical images, which typically exhibit extremely sparse visual evidence to inform clinical decision-making. We recognize that pruning visual tokens outside the grounding region greatly enhances medical reasoning. However, a united RL framework for active visual token pruning (VTP) and medical multimodal reasoning remains unestablished. Here, we propose a dual-stream RL framework, ViToS, to fulfill token pruning and question answering. ViToS trains one policy model with two task branches, where one focuses on grounding while the other conducts token-sparse reasoning after VTP. Furthermore, we solve the coupled policy learning problem by introducing the cross-feedback sequential optimization, avoiding gradient conflict and facilitating convergence of the shared policy model. Evaluated on seven medical benchmarks, our method reduces visual tokens to 77% of the original sequence length while achieving a 108.27% relative performance on Lingshu-7B and 104.16% relative performance on HuatuoGPT-Vision-7B. Overall, ViToS delivers superior performance and inference speedup, establishing an efficient paradigm for medical multimodal reasoning.
Kaitao Chen, Weiqian Zhao, Jiamin Wu +6
1Fudan University · 2Shanghai Artificial Intelligence Laboratory · 3Shanghai Jiao Tong University +1
Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain predominantly outcome-centric, relying on final answer correctness or sequence-level preferences. This suffers from sparse credit assignment, making it difficult to optimize the reasoning process essential for clinical applications. Our analysis reveals that cascading errors from early-stage reasoning failures are a leading cause of incorrect predictions in medical visual question answering (VQA) benchmarks. Motivated by this, we propose Medical Reasoning-aware Policy Optimization (MRPO), an RL algorithm that incorporates step-wise process rewards. When the final answer is incorrect, MRPO assigns exponentially larger penalties to tokens in earlier invalid reasoning steps, breaking failure cascades without compromising successful paths. Across four multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3-VL-8B-Thinking even surpasses substantially larger medical MLLMs such as HuatuoGPT-Vision-34B by 4.59 points. Moreover, MRPO reduces early-stage reasoning failures from 58.6% to 13.4%, showing that targeted mitigation of cascading failures improves both reasoning quality and final answer accuracy. Our code is available at https://github.com/dmis-lab/MRPO
Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents must aggregate evidence across ordered images, remain robust to answer-order perturbations, and avoid overfitting to noisy search-time feedback. We study MedFrameQA through a controlled comparison of five inference-time agentic strategies, optimized using the same high-budget ShinkaEvolve configuration and evaluated on a reproducible internal frozen split (1,331 evolution, 665 holdout, 855 final test). Across five independent repeated runs, the strongest method emerges as the simplest robust aggregator: the \textbf{order-vote} policy achieves 57.89±0.65% final-test accuracy, significantly outperforming the fixed baseline (52.73±0.42%) and the more complex, albeit brittle, order-rerank variant (55.79±0.43%). Paired bootstrap analysis confirms these significant gains. Extending the evolutionary search budget from 50 to 100 generations yields no generalization benefit: while holdout performance marginally increases, final-test accuracy drops from 57.89% to 56.02%. Our findings suggest that for multi-image medical reasoning, defining the correct agentic decision rule is substantially more impactful than expanding the optimization search budget.