Multi-agent debate (MAD) is often used to improve large language model (LLM) reasoning, but sequential debate is rarely a neutral aggregator of agents' opinions. We show that sequential MAD suffers from a pronounced first-speaker bias: agents disproportionately shape the final answer when they speak first. As a result, placing a stronger model after weaker ones can substantially offset its reasoning advantage. We then focus on the disadvantaged strong-agent-last setting and ask whether personality prompting can mitigate this imbalance. Drawing on the Big Five model, we study agreeableness and extraversion as behavioral interventions applied to either the strong or weak side. We find that their effects are trait-specific. Influence consistently shifts in the direction of lower agreeableness, and assigning low agreeableness to the stronger agent helps restore its lost influence and improves final accuracy. Extraversion, by contrast, produces less systematic changes in influence and accuracy, with its clearest effect appearing in agents' verbosity. These findings show that effective MAD design depends not only on model capability, but also on how speaking order and induced interaction behavior shape the debate process.
Figures & tables
Figure 1: Results under the no-personality baseline, shown by benchmark dataset and averaged across model combinations. Left: Mean agent influence across debate orders, with stacked segments showing the relative influence of the strong and weak agents within each order. Right: Mean final debate accuracy across debate orders.
Weak
Strong
W-W-S
S-W-W
Acc.
41.29
62.32
66.43
68.78
Infl.
N.A.
N.A.
58.83
79.84
Table 1: Average standalone and debate accuracy (%), and strong-agent influence (%) for the no-personality debate baselines, aggregated across benchmark datasets and model combinations.
Strong-Agent Influence
Final Accuracy
Statistic
ah_ah_n
al_al_n
n_n_ah
n_n_al
ah_ah_n
al_al_n
n_n_ah
n_n_al
Mean Δ (pp)
3.76
-4.84
-11.07
9.55
-0.02
-2.42
-5.13
1.64
Std
5.33
9.77
8.13
6.89
2.49
4.11
2.86
3.42
Paired t-test p
<0.001 ∗∗∗
0.0033 ∗∗
<0.001 ∗∗∗
<0.001 ∗∗∗
0.5231
0.9997
1.0000
0.0021 ∗∗
Wilcoxon p
<0.001 ∗∗∗
0.0120 ∗
<0.001 ∗∗∗
<0.001 ∗∗∗
0.6171
0.9993
1.0000
0.0042 ∗∗
Sign test p
0.0022 ∗∗
0.2682
<0.001 ∗∗∗
<0.001 ∗∗∗
0.3136
0.9996
1.0000
0.0069 ∗∗
Table 2: Percentage point differences in strong-agent influence and final debate accuracy under agreeableness manipulations relative to the no-personality baseline in the Weak-Weak-Strong order, aggregated across benchmark datasets and model combinations. Accuracy tests are one-tailed, influence tests are two-tailed. Significance levels: ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 , †p<0.1 .
Strong-Agent Influence
Final Accuracy
Statistic
eh_eh_n
el_el_n
n_n_eh
n_n_el
eh_eh_n
el_el_n
n_n_eh
n_n_el
Mean Δ (pp)
3.92
3.94
-3.04
-4.14
0.09
0.78
-2.21
-2.53
Std
6.56
5.09
5.40
6.32
3.22
2.67
2.51
2.31
Paired t-test p
<0.001 ∗∗∗
<0.001 ∗∗∗
<0.001 ∗∗∗
<0.001 ∗∗∗
0.4320
0.0365 ∗
1.0000
1.0000
Wilcoxon p
<0.001 ∗∗∗
<0.001 ∗∗∗
0.0014 ∗∗
<0.001 ∗∗∗
0.3500
0.0379 ∗
1.0000
1.0000
Sign test p
0.0022 ∗∗
0.0064 ∗∗
0.0166 ∗
0.0064 ∗∗
0.5660
0.0939 †
1.0000
1.0000
Table 3: Percentage point differences in strong-agent influence and final debate accuracy under extraversion manipulations relative to the no-personality baseline in the Weak-Weak-Strong order, aggregated across benchmark datasets and model combinations. Accuracy tests are one-tailed, influence tests are two-tailed. Significance levels: ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 , †p<0.1 .
Figure 2: Median justification length by extraversion configuration and agent position in the W-W-S order, aggregated across benchmark datasets and model combinations. Bold labels indicate the strong agent. Connected annotations show the percentage change in justification length between the high- and low-extraversion conditions, computed as (low−high)/high×100% .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Weak Model
Strong Model
gemini-2.5-flash-lite
gemini-2.5-flash
gemini-2.5-flash-lite
gemini-3-flash-preview
gpt-4o-mini
gemini-2.5-flash
gpt-4o-mini
gemini-3-flash-preview
llama-4-scout
gemini-2.5-flash
llama-4-scout
gemini-3-flash-preview
Appendix
Table 4: Model combinations used in debate experiments.
Agreeableness
Extraversion
Model
No-personality
High
Low
High
Low
2.5 Flash
52.60
52.10
53.48
51.97
53.48
2.5 Flash Lite
44.35
43.23
42.97
42.22
44.36
3 Flash
72.04
69.66
70.79
70.91
71.04
4o Mini
36.46
36.46
36.22
36.21
36.71
LLaMA 4 Scout
40.72
39.47
40.22
39.59
39.46
Appendix
Table 5: Single-agent accuracy (%) under no-personality and personality-prompted conditions, averaged across benchmark datasets.
Statistic
Accuracy
Strong-agent influence
Mean Δ (pp)
2.34
21.01
Std.
2.94
13.22
Paired t-test p
<0.001∗∗∗
<0.001∗∗∗
Wilcoxon p
<0.001∗∗∗
<0.001∗∗∗
Sign test p
<0.001∗∗∗
<0.001∗∗∗
Mixed-effects p
<0.001∗∗∗
<0.001∗∗∗
Appendix
Table 6: Paired comparison of the S-W-W and W-W-S debate orders. Values are computed as S-W-W minus W-W-S for final accuracy and strong-agent influence, measured in percentage points. Reported p -values are one-tailed for accuracy and two-tailed for strong-agent influence. Significance levels: ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 .
Effect
Strong-Agent Influence
Accuracy
S-W-W
21.01∗∗∗
2.34∗
Capability Gap
0.90∗∗∗
0.37∗∗∗
S-W-W × Capability Gap
−0.39∗
0.05
Appendix
Table 7: Mixed-effects analysis of whether the capability gap between the strong and weak agents moderates the effect of speaking order. Capability gap is defined as the difference in no-personality standalone accuracy between the strong and weak models and is mean-centered before estimation. W-W-S is the reference speaking order. Coefficients are measured in percentage points. Significance levels: ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 .
Weak-Weak-Strong
Statistic
no-personality
n_n_al
ah_ah_al
Accuracy mean
66.43
68.08
67.23
Accuracy std.
14.93
14.16
14.81
Strong-agent influence mean
58.83
68.38
71.43
Strong-agent influence std.
17.92
15.69
15.06
Appendix
Table 8: Final debate accuracy (%) and strong-agent influence (%) for selected agreeableness combinations in the Weak-Weak-Strong order, aggregated across benchmark datasets and model combinations.
Figure 3: Median justification length by agreeableness configuration and agent position in the W-W-S order, aggregated across benchmark datasets and model combinations. Bold labels indicate the strong agent. Connected annotations show the percentage change in justification length between the low- and high-agreeableness variants, computed as (high−low)/low×100% .
Agreeableness
Extraversion
Statistic
ah_ah_n
al_al_n
n_n_ah
n_n_al
eh_eh_n
el_el_n
n_n_eh
n_n_el
Accuracy (%)
72.67
70.67
71.29
74.17
72.67
72.17
70.79
73.30
Strong-agent influence (%)
81.54
78.13
78.56
82.71
79.81
79.67
77.51
79.62
Appendix
Table 9: Final debate accuracy and strong-agent influence under the alternative personality prompt in the Weak-Weak-Strong order for the GPT-4o-mini/Gemini-3-Flash model combination, aggregated across the four benchmark datasets.
Figure 4: Pilot analysis across debate rounds, averaged over 50 randomly sampled GPQA-Diamond questions and model combinations. Left: Majority-answer accuracy across debate rounds. Right: Inter-agent unanimous agreement rate across debate rounds.
Condition
Agent
Answer
Round-1 justification
n_n_n
Weak 1
H
Key dimensions, particularly thickness and width, are standardized to ensure proper fit and torque transmission.
n_n_n
Weak 2
H
Key dimensions such as thickness and width are standardized to ensure a precise fit within the keyway.
n_n_n
Strong
H
Key profile dimensions, including thickness, are standardized to ensure proper fit and torque transmission between a shaft and hub.
n_n_al
Weak 1
H
Key dimensions, particularly thickness and width, are standardized to ensure proper fit and torque transmission.
n_n_al
Weak 2
H
Key dimensions such as thickness and width are standardized to ensure a precise fit within the keyway.
n_n_al
Strong
B
The diameter of the shaft is the primary criterion for selecting key profile dimensions from standards. The other agents’ focus on key thickness is a consequence, not the initial determinant.
Appendix
Table 10: Representative first-round trace for SuperGPQA Question 9 under the W-W-S order. Response text is shortened for presentation.
Multi-Agent Debate (MAD) improves the reasoning performance of Large Language Models (LLMs) through multi-round interaction. However, LLMs in MAD are highly susceptible to blind conformity. Existing individual evaluation methods, typically based on confidence or perplexity, fail to reflect the correctness of reasoning and may even exacerbate blind conformity. To address this, we shift the perspective from individual evaluation to group interaction. We define mutual referencing among LLMs as \textbf{Debate Relationships} and recognize that regulating these relationships is the key to mitigating blind conformity. In this paper, we propose a novel framework for \textbf{D}ynamically r\textbf{E}gulating deb\textbf{A}te \textbf{R}elationships (DEAR) from the group perspective. At first, DEAR quantifies consensus and divergence as \textit{group evidence} to capture the debate state. Then, DEAR operates through three stages: 1) What: perceiving group consultation tendency and uncertainty; 2) Who: introducing a Selection RL-Agent to dynamically select reference peers; and 3) How: adopting a Behavior RL-Agent to adaptively adjust generation behaviors. Notably, we formulate the execution of the two RL-Agents as a sequential decision-making process, jointly optimizing via multi-agent reinforcement learning. Extensive experiments demonstrate that DEAR achieves superior performance while significantly reducing token consumption.
Hao Wu, Shoucheng Song, Chang Yao +4
School of Computer Science & Technology, Beijing Jiaotong University, Beijing, China · Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing, China
Multi-agent debate (MAD) can improve large language model reasoning, but fixed debate pipelines often waste computation and can amplify correlated errors among similar agents. We propose ARMOR-MAD, a training-free heterogeneous MAD framework that treats debate as conditional computation. ARMOR-MAD combines three components: Pre-debate Agreement Routing (PAR) decides whether independently generated Round-0 answers require debate; Early Agreement Stopping Evaluator (EASE) stops debate after convergence; and Semantic Outlier Detection (SOD) down-weights abnormal final answers during aggregation. Across MATH Level 5, GSM8K, MMLU, and MMLU-Pro, ARMOR-MAD consistently improves over fixed-round heterogeneous debate with the same model pool, reaching 65.5%, 96.5%, 90.0%, and 81.5% accuracy, respectively. The results suggest that genuine model heterogeneity and agreement-based control are both important for making MAD more accurate and efficient.
Fuqiang Niu, Bowen Zhang
School of Cyber Science and Technology, University of Science and Technology of China, Hefei, China · School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the agents' inherently biased concept priors unaddressed. To mitigate this systematic weakness, we propose R2-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates. R2-MAD intervenes on both failure modes through two complementary mechanisms: A debate-state-aware retrieval policy dynamically calibrates the concept prior by retrieving relevant historical evidence based on the current consensus level. Then these retrieved experiences provide a basis for estimating per-agent reliability, yielding confidence weights to modulate peer influence. Experiments on various benchmarks show that R2-MAD achieves consistent improvements over existing single-agent and MAD baselines.
Xuanfa Jin, Zhijian Ma, Yongcheng Zeng +3
Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University College London