Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variation in the visual scale needed across samples and agents. In addition, these frameworks frequently suffer from groupthink, a phenomenon where agents prematurely abandon correct deductions to conform with confident but hallucinated peer responses. To address these bottlenecks, we introduce DREAM (Dynamic Resolution Assignment For Multimodal Multi-Agent Debate), which operates via two core components: (1) Dynamic Resolution Assignment, a zero-shot probe round where agents test multiple resolutions, quantify uncertainty using Average Normalized Log-Likelihood (ANLL), and use an adaptive threshold to assign each agent to its empirically optimal resolution; (2) Uncertainty-Guided Rollback Aggregation counters groupthink by tracking each agent's uncertainty over rounds and restoring early low-uncertainty answers overridden by group pressure. On six multimodal datasets, DREAM improves the accuracy-token trade-off over multi-agent debate baselines by 1.5-3.2% accuracy without dataset-specific tuning.
Figures & tables
Figure 1: Visual tasks demand different resolutions: simple perception tasks (left) only need 0.5× , while 2.0× introducing long-context noise that actively degrades accuracy. Conversely, dense charts (right) require 2.0× to preserve critical fine-grained details, thus 0.5× irreversibly discarding them. Standard MAD uses a fixed resolution, degrading on simple samples and failing on complex ones. DREAM solves this by dynamically assigning each sample to its empirically optimal scale.
Figure 2: The DREAM framework pipeline. Top: DRA’s probe-round decision tree, selecting a DRA strategy based first on confident-set size ∣C∣ and then, when ∣C∣≥2 , on answer agreement within C ; the disagreement strategy ( ConflictMixRes ) is highlighted in orange. Bottom: URA’s per-agent ANLL trajectory over debate rounds r=1,2,…,R , comparing each agent’s final round to its earlier peak r∗ (gold marker) against the rescue threshold δi . Agent 1 (pink) exceeds δ1 and is rescued; Agents 2–3 (green) keep their final answers. Green marks the default path throughout.
Model
Method
Know.
Mathematics
Med.
Reasoning
Avg.
Δ
MME
MVt
MVs
PVQA
MME-R
VP
Gemma-3-4B
IO
79.3
37.4
15.0
24.9
21.8
22.9
33.5
CoT [ 33 ]
80.2
37.4
15.7
24.9
21.8
22.9
33.8
SC [ 32 ]
80.1
37.1
15.1
24.9
20.8
22.8
33.5
DwI [ 7 ]
58.9
23.7
8.7
26.1
23.0
18.6
26.5
LLM Debate [ 6 ]
78.4
57.0
24.3
40.6
26.5
36.3
43.9
Table 1: Performance on M3MAD-Bench across four distinct model families (results for an additional model are provided in Appendix Table 5 ). Best results are bolded , second-best underlined . Δ denotes DREAM’s average improvement over the LLM Debate baseline. MVt : MathVista, MVs : MathVision, PVQA : PathVQA, MME-R : MME-Reasoning, VP : VisualPuzzles.
Model
Method
Know.
Mathematics
Med.
Reasoning
Avg.
Δ
MME
MVt
MVs
PVQA
MME-R
VP
Gemma-3-4B
Perceiver [ 11 ]
w/ LLM Debate [ 6 ]
82.5
52.4
23.1
35.0
26.0
32.7
41.9
w/ Div-MAD [ 17 ]
43.0
45.5
20.6
18.3
20.5
30.3
29.7
w/ DMAD [ 19 ]
81.7
56.1
22.1
38.9
25.8
30.0
42.4
DREAM
Table 2: Performance of the evaluated models under different wrapper methods ( Perceiver vs. DREAM ). Best results are bolded , second-best underlined . Δ denotes the average accuracy improvement of DREAM over the corresponding bare debate baseline. MVt : MathVista, MVs : MathVision, PVQA : PathVQA, MME-R : MME-Reasoning, VP : VisualPuzzles.
Method
Gemma-3-4B
Qwen2.5-VL-7B
InternVL3-8B
Baseline (LLM Debate)
43.9
47.0
48.9
+ DRA (Resolution Probing)
44.7 ( ↑ 0.8 )
48.3 ( ↑ 1.3 )
49.6 ( ↑ 0.7 )
+ URA (Adaptive Threshold)
45.1 ( ↑ 1.2 )
48.9 ( ↑ 1.9 )
50.5 ( ↑ 1.6 )
+ URA (Rollback)
45.4 ( ↑ 1.5 )
50.2 ( ↑ 3.2 )
51.4 ( ↑ 2.5 )
Table 3: Ablation of DREAM components and their contribution to average accuracy (%).
Figure 3: Accuracy vs. prompt token consumption for Qwen2.5-VL-7B between baselines and DREAM.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Prompt Tokens (M)
Accuracy (%)
Div-MAD
2.57
36.8
w/ DREAM
4.07
38.1
LLM Debate
9.38
47.0
w/ DREAM
12.18
50.2
DMAD
27.96
47.3
w/ DREAM
30.88
49.5
Appendix
Table 4: Prompt token consumption and accuracy for Qwen2.5-VL-7B.
Model
Method
Know.
Mathematics
Med.
Reasoning
Avg.
Δ
MME
MVt
MVs
PVQA
MME-R
VP
LLaVA-Next-7B
IO
40.2
24.0
8.0
11.2
12.8
18.0
19.0
CoT [ 33 ]
30.2
24.0
7.2
4.6
14.8
16.0
16.1
SC [ 32 ]
29.2
24.8
9.8
3.4
13.4
20.0
16.8
DwI [ 7 ]
58.9
27.2
9.3
29.5
16.0
20.8
26.9
LLM Debate [ 6 ]
72.0
34.5
9.9
23.7
17.3
22.4
30.0
Appendix
Table 5: Performance on M3MAD-Bench multimodal benchmarks for LLaVA-Next-7B. Best result per column among rows with reported values bolded , second-best underlined . Δ denotes w/ DREAM ’s average accuracy change over its bare debate baseline. MVt : MathVista, MVs : MathVision, PVQA : PathVQA, MME-R : MME-Reasoning, VP : VisualPuzzles.
Resolution Strategy
InternVL3-14B
InternVL3-8B
Qwen2.5-VL-7B
Average
0.5 × Scale
49.84
49.93
49.45
49.74
2 × Scale
49.78
49.48
49.31
49.52
4 × Scale
50.24
50.35 †
49.23
49.94 †
Appendix
Table 6: Resolution scale ablation (Accuracy %). † 4 × Scale is not stable.
Threshold
InternVL3-14B
InternVL3-8B
Qwen2.5-VL-7B
Avg.
Δ
Static 0.10
50.25
49.66
49.29
49.73
-0.01
Static 0.20
50.12
50.16
48.93
49.74
0.00
Static 0.30
50.51
49.57
48.90
49.66
↓ 0.08
Static 0.50
50.36
49.60
49.05
49.67
↓ 0.07
Adaptive ( μ−σ )
50.70
50.40
49.50
50.20
↑ 0.46
Appendix
Table 7: Ablation on uncertainty thresholds across models (Accuracy %). Δ is each row’s average minus the best static-threshold average in hindsight (49.74, achieved at 0.20).
Figure 4: Share of gated samples routed to each DRA branch, by benchmark. ConflictMixRes dominates on MathVision (64%), MME-Reasoning (41%), and VisualPuzzles (41%), while MixRes dominates on PathVQA (67%); StaticAll is comparatively more common on MME (34%) and MathVista (28%). No single branch dominates uniformly across benchmarks, consistent with the probe resolving a genuinely per-sample, per-dataset routing decision rather than converging to one default behavior.
Figure 5: Same branch-share data as Figure 4 , shown as a benchmark-by-branch heatmap for direct cross-dataset comparison.
Figure 6: Accuracy (left) and mean token cost (right) per DRA branch, pooled across all six benchmarks, against the DRA-average and the no-DRA baseline. StaticAll is the most accurate individual branch (66.6%) but also among the most expensive (35,846 tokens/sample); ConflictMixRes is the least accurate (37.5%) despite a high token cost (33,903), reflecting that it fires precisely when confident agents disagree and no majority resolution can be trusted. The DRA-average (50.2%, 30,989 tokens) exceeds the no-DRA baseline (47.0%, 21,416 tokens) in accuracy at roughly 1.45× the token cost.
Figure 7: Distribution of total tokens per sample for each DRA branch, split by whether the final answer was correct. Within most branches, incorrect samples show a higher median and a longer upper tail than correct samples, most visibly for StaticAll , whose incorrect-sample whisker extends past 165k tokens; the no-DRA baseline shows the same correct/incorrect gap at a lower absolute cost.
Figure 8: Mitigating Visual Hallucinations. Qualitative comparison on the MME dataset. While the standard baseline enters a compounding hallucination loop (detailing non-existent flowers and bumblebees), DREAM’s DRA module detects high confidence at the 2.0× resolution and routes the debate optimally, yielding the correct answer with high precision.
Figure 9: Dismantling Groupthink. Qualitative comparison on MathVision. The standard baseline succumbs to agreement bias, where subsequent agents blindly agree with an initial flawed assumption. DREAM detects early disagreement during probing and deploys a mixed-resolution payload, forcing the agents to successfully ground their spatial reasoning rather than relying on unverified assumptions.
Figure 10: Robust Clinical Synthesis. Qualitative comparison on PathVQA. While the fixed-resolution baseline struggles to extract specific etiology (providing a generic ”third-degree” diagnosis), DREAM’s multi-scale probing seamlessly extracts complementary clinical cues across distinct visual resolutions. The resulting mixed-resolution debate yields a far more precise and comprehensive medical evaluation.
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the agents' inherently biased concept priors unaddressed. To mitigate this systematic weakness, we propose R2-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates. R2-MAD intervenes on both failure modes through two complementary mechanisms: A debate-state-aware retrieval policy dynamically calibrates the concept prior by retrieving relevant historical evidence based on the current consensus level. Then these retrieved experiences provide a basis for estimating per-agent reliability, yielding confidence weights to modulate peer influence. Experiments on various benchmarks show that R2-MAD achieves consistent improvements over existing single-agent and MAD baselines.
Xuanfa Jin, Zhijian Ma, Yongcheng Zeng +3
Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University College London
Multi-Agent Debate (MAD) improves the reasoning performance of Large Language Models (LLMs) through multi-round interaction. However, LLMs in MAD are highly susceptible to blind conformity. Existing individual evaluation methods, typically based on confidence or perplexity, fail to reflect the correctness of reasoning and may even exacerbate blind conformity. To address this, we shift the perspective from individual evaluation to group interaction. We define mutual referencing among LLMs as \textbf{Debate Relationships} and recognize that regulating these relationships is the key to mitigating blind conformity. In this paper, we propose a novel framework for \textbf{D}ynamically r\textbf{E}gulating deb\textbf{A}te \textbf{R}elationships (DEAR) from the group perspective. At first, DEAR quantifies consensus and divergence as \textit{group evidence} to capture the debate state. Then, DEAR operates through three stages: 1) What: perceiving group consultation tendency and uncertainty; 2) Who: introducing a Selection RL-Agent to dynamically select reference peers; and 3) How: adopting a Behavior RL-Agent to adaptively adjust generation behaviors. Notably, we formulate the execution of the two RL-Agents as a sequential decision-making process, jointly optimizing via multi-agent reinforcement learning. Extensive experiments demonstrate that DEAR achieves superior performance while significantly reducing token consumption.
Hao Wu, Shoucheng Song, Chang Yao +4
School of Computer Science & Technology, Beijing Jiaotong University, Beijing, China · Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing, China
Multi-agent debate (MAD) can improve large language model reasoning, but fixed debate pipelines often waste computation and can amplify correlated errors among similar agents. We propose ARMOR-MAD, a training-free heterogeneous MAD framework that treats debate as conditional computation. ARMOR-MAD combines three components: Pre-debate Agreement Routing (PAR) decides whether independently generated Round-0 answers require debate; Early Agreement Stopping Evaluator (EASE) stops debate after convergence; and Semantic Outlier Detection (SOD) down-weights abnormal final answers during aggregation. Across MATH Level 5, GSM8K, MMLU, and MMLU-Pro, ARMOR-MAD consistently improves over fixed-round heterogeneous debate with the same model pool, reaching 65.5%, 96.5%, 90.0%, and 81.5% accuracy, respectively. The results suggest that genuine model heterogeneity and agreement-based control are both important for making MAD more accurate and efficient.
Fuqiang Niu, Bowen Zhang
School of Cyber Science and Technology, University of Science and Technology of China, Hefei, China · School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China