When Order Matters: First-Speaker Bias and Mitigation through Personality in Sequential Multi-Agent Debate
Organizations: School of Computing National University of Singapore Singapore
Abstract
Multi-agent debate (MAD) is often used to improve large language model (LLM) reasoning, but sequential debate is rarely a neutral aggregator of agents' opinions. We show that sequential MAD suffers from a pronounced first-speaker bias: agents disproportionately shape the final answer when they speak first. As a result, placing a stronger model after weaker ones can substantially offset its reasoning advantage. We then focus on the disadvantaged strong-agent-last setting and ask whether personality prompting can mitigate this imbalance. Drawing on the Big Five model, we study agreeableness and extraversion as behavioral interventions applied to either the strong or weak side. We find that their effects are trait-specific. Influence consistently shifts in the direction of lower agreeableness, and assigning low agreeableness to the stronger agent helps restore its lost influence and improves final accuracy. Extraversion, by contrast, produces less systematic changes in influence and accuracy, with its clearest effect appearing in agents' verbosity. These findings show that effective MAD design depends not only on model capability, but also on how speaking order and induced interaction behavior shape the debate process.
Figures & tables
| Weak | Strong | W-W-S | S-W-W | |
|---|---|---|---|---|
| Acc. | 41.29 | 62.32 | 66.43 | 68.78 |
| Infl. | N.A. | N.A. | 58.83 | 79.84 |
| Strong-Agent Influence | Final Accuracy | |||||||
|---|---|---|---|---|---|---|---|---|
| Statistic | ah_ah_n | al_al_n | n_n_ah | n_n_al | ah_ah_n | al_al_n | n_n_ah | n_n_al |
| Mean (pp) | 3.76 | -4.84 | -11.07 | 9.55 | -0.02 | -2.42 | -5.13 | 1.64 |
| Std | 5.33 | 9.77 | 8.13 | 6.89 | 2.49 | 4.11 | 2.86 | 3.42 |
| Paired t-test | <0.001 ∗∗∗ | 0.0033 ∗∗ | <0.001 ∗∗∗ | <0.001 ∗∗∗ | 0.5231 | 0.9997 | 1.0000 | 0.0021 ∗∗ |
| Wilcoxon | <0.001 ∗∗∗ | 0.0120 ∗ | <0.001 ∗∗∗ | <0.001 ∗∗∗ | 0.6171 | 0.9993 | 1.0000 | 0.0042 ∗∗ |
| Sign test | 0.0022 ∗∗ | 0.2682 | <0.001 ∗∗∗ | <0.001 ∗∗∗ | 0.3136 | 0.9996 | 1.0000 | 0.0069 ∗∗ |
| Strong-Agent Influence | Final Accuracy | |||||||
|---|---|---|---|---|---|---|---|---|
| Statistic | eh_eh_n | el_el_n | n_n_eh | n_n_el | eh_eh_n | el_el_n | n_n_eh | n_n_el |
| Mean (pp) | 3.92 | 3.94 | -3.04 | -4.14 | 0.09 | 0.78 | -2.21 | -2.53 |
| Std | 6.56 | 5.09 | 5.40 | 6.32 | 3.22 | 2.67 | 2.51 | 2.31 |
| Paired t-test | <0.001 ∗∗∗ | <0.001 ∗∗∗ | <0.001 ∗∗∗ | <0.001 ∗∗∗ | 0.4320 | 0.0365 ∗ | 1.0000 | 1.0000 |
| Wilcoxon | <0.001 ∗∗∗ | <0.001 ∗∗∗ | 0.0014 ∗∗ | <0.001 ∗∗∗ | 0.3500 | 0.0379 ∗ | 1.0000 | 1.0000 |
| Sign test | 0.0022 ∗∗ | 0.0064 ∗∗ | 0.0166 ∗ | 0.0064 ∗∗ | 0.5660 | 0.0939 † | 1.0000 | 1.0000 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Weak Model | Strong Model |
|---|---|
| gemini-2.5-flash-lite | gemini-2.5-flash |
| gemini-2.5-flash-lite | gemini-3-flash-preview |
| gpt-4o-mini | gemini-2.5-flash |
| gpt-4o-mini | gemini-3-flash-preview |
| llama-4-scout | gemini-2.5-flash |
| llama-4-scout | gemini-3-flash-preview |
| Agreeableness | Extraversion | ||||
|---|---|---|---|---|---|
| Model | No-personality | High | Low | High | Low |
| 2.5 Flash | 52.60 | 52.10 | 53.48 | 51.97 | 53.48 |
| 2.5 Flash Lite | 44.35 | 43.23 | 42.97 | 42.22 | 44.36 |
| 3 Flash | 72.04 | 69.66 | 70.79 | 70.91 | 71.04 |
| 4o Mini | 36.46 | 36.46 | 36.22 | 36.21 | 36.71 |
| LLaMA 4 Scout | 40.72 | 39.47 | 40.22 | 39.59 | 39.46 |
| Statistic | Accuracy | Strong-agent influence |
|---|---|---|
| Mean (pp) | 2.34 | 21.01 |
| Std. | 2.94 | 13.22 |
| Paired t-test | ||
| Wilcoxon | ||
| Sign test | ||
| Mixed-effects |
| Effect | Strong-Agent Influence | Accuracy |
|---|---|---|
| S-W-W | ||
| Capability Gap | ||
| S-W-W Capability Gap |
| Weak-Weak-Strong | |||
|---|---|---|---|
| Statistic | no-personality | n_n_al | ah_ah_al |
| Accuracy mean | 66.43 | 68.08 | 67.23 |
| Accuracy std. | 14.93 | 14.16 | 14.81 |
| Strong-agent influence mean | 58.83 | 68.38 | 71.43 |
| Strong-agent influence std. | 17.92 | 15.69 | 15.06 |
| Agreeableness | Extraversion | |||||||
|---|---|---|---|---|---|---|---|---|
| Statistic | ah_ah_n | al_al_n | n_n_ah | n_n_al | eh_eh_n | el_el_n | n_n_eh | n_n_el |
| Accuracy (%) | 72.67 | 70.67 | 71.29 | 74.17 | 72.67 | 72.17 | 70.79 | 73.30 |
| Strong-agent influence (%) | 81.54 | 78.13 | 78.56 | 82.71 | 79.81 | 79.67 | 77.51 | 79.62 |
| Condition | Agent | Answer | Round-1 justification |
|---|---|---|---|
| n_n_n | Weak 1 | H | Key dimensions, particularly thickness and width, are standardized to ensure proper fit and torque transmission. |
| n_n_n | Weak 2 | H | Key dimensions such as thickness and width are standardized to ensure a precise fit within the keyway. |
| n_n_n | Strong | H | Key profile dimensions, including thickness, are standardized to ensure proper fit and torque transmission between a shaft and hub. |
| n_n_al | Weak 1 | H | Key dimensions, particularly thickness and width, are standardized to ensure proper fit and torque transmission. |
| n_n_al | Weak 2 | H | Key dimensions such as thickness and width are standardized to ensure a precise fit within the keyway. |
| n_n_al | Strong | B | The diameter of the shaft is the primary criterion for selecting key profile dimensions from standards. The other agents’ focus on key thickness is a consequence, not the initial determinant. |