Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group.
Figures & tables
Figure 1: Illustration of the contaminated multi-agent deliberation setup used throughout the experiments. Red robots are deceptive agents . Green robots are honest agents. Agents deliberate sequentially in public and reflect privately, with honest agents working to solve the question and deceptive agents attempting to steer the group toward an incorrect conclusion. At the end of voting, some agents will be swayed towards defecting to an incorrect answer, like the agent in the last row whose text bubble is highlighted in red . In this work, we aim to quantify how this effect scales with respect to the prevalence of deceptive agents in parties of varying sizes.
Figure 2: Example of coordination among three deceivers.
Honest agents
k/N=0
k/N=1/5
k/N=1/3
k/N=3/7
2
2+0 (2)
–
2+1 (3)
–
4
4+0 (4)
–
4+2 (6)
4+3 (7)
8
8+0 (8)
8+2 (10)
8+4 (12)
8+6 (14)
12
–
12+3 (15)
12+6 (18)
12+9 (21)
Table 1: We test the following group compositions, shown as honest agents + deceivers (total group size N in parentheses). This lets us compare groups with the same deceiver proportion or the same number of deceivers.
R2
Model
Linear slope p
Straight line
Square root
Step
Gemini 3.8 Flash
<0.001
0.97
0.84
0.54
Grok 4.3
<0.001
0.90
0.88
0.72
DeepSeek V4.1 Flash
<0.01
0.82
0.73
0.52
Muse Glimmer
<0.001
0.97
0.94
0.74
Table 2: Relationship between adversarial proportion and honest defection. R2 compares linear, square-root, and step fits, weighted by the number of initially correct honest agents; the straight line fits best for every model.
Figure 3: Honest defection by adversarial proportion, the number of deceivers divided by the number of agents in the group ( k/N ), pooled across group sizes (0 marks the baseline without deceivers). Defection increases approximately linearly with adversarial proportion in every model. Error bars show 95% percentile bootstrap confidence intervals.
Figure 4: Honest defection in Gemini 3.8 Flash by group size, with one panel per adversarial proportion. Across all models, larger groups do not consistently show lower defection at a fixed proportion, and group size magnitude visibly affects honest defection less than adversarial proportion. Error bars show 95% percentile bootstrap confidence intervals.
Figure 5: Honest defection with non-coordinated and coordinated deceivers. Defection is lower under coordination at each nonzero proportion, and the trends in both conditions also have linear best-fits. Error bars show 95% percentile bootstrap confidence intervals.
Figure 6: Honest defection for four honest–deceiver model pairings at two adversarial proportions. Muse Glimmer, the more sycophantic honest model, defects more often than Gemini. DeepSeek, the more persuasive deceiver, generally causes more defection than Grok. The identity of the honest side has more impact on defection rate than that of the deceiver side. Error bars show 95% percentile bootstrap confidence intervals.
Figure 7: Distribution of most common persuasion tactics used by deceivers. (a) Each model’s three most prevalent tactics across rounds 0 and 1. (b) Prevalence by round across models. Multiple tactics may appear in the same message. Definitions are in Appendix D .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Model
Questions ( 1/4 , 2/4 , 3/4 )
Trials
Gemini 3.8 Flash
107 (42, 30, 35)
1,284
Grok 4.3
95 (38, 30, 27)
1,140
DeepSeek V4.1 Flash
95 (38, 30, 27)
1,140
Muse Glimmer
102 (50, 32, 20)
1,224
Appendix
Table 3: Questions and trials in the main experiments, grouped by model. Parentheses give the numbers of questions answered correctly once, twice, or three times in four independent attempts.
The effectiveness of multi-agent LLM deliberation depends not only on the agents' individual predictions, but also on how they communicate and collaborate. We study this mechanism through the lens of Friedkin-Johnsen (FJ) opinion dynamics, a tractable model for analyzing stubbornness, influence, and opinion change in multi-agent systems that captures empirically observed deliberation patterns. We show that the FJ parameters are input-dependent, turning multi-agent deliberation into a mixture of experts. This perspective implies that multi-agent systems can outperform single agents and static ensembles when routing reflects agent competence. Since competence is latent in practice, we analyze how influence is established through observable proxies: agents' self-assessed confidence, their perceived confidence, and initial alignment with other agents' views.
Franka Bause, Jonas Niederle, Martin Pawelczyk +1
CISPA Helmholtz Center for Information Security, Saarbrücken, Germany · Faculty of Computer Science, University of Vienna, Vienna, Austria.
Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that convergence reflects genuine deliberation or social compliance. We show that the conventional answer flip rate conflates three distinct mechanisms: spontaneous instability, stance-induced conformity, and reasoning-induced persuasion. Our three-source decomposition framework isolates each through controlled counterfactual conditions. In the primary MMLU-Pro setting, 37% of agent-question observations change under self-reflection alone, while robustness tests show substantial model-dependent instability across GPQA-Diamond and three model families; strict conformity is 29% in the primary setting and remains predominantly harmful across model replications (57-77% correct-to-wrong). A controlled information-gradient experiment reveals that even vacuous reasoning is associated with 20-39% error adoption among resistant agents, with reasoning-like presentation carrying substantial persuasive weight. Harmful conformity can be predicted from Round 0 features (AUC = 0.79), and risk-targeted intervention reduces it by 13.6 percentage points (p < 0.001). However, without correctness labels or self-reflection controls, reducing peer adoption does not improve accuracy, because harmful and beneficial influence cannot be distinguished.
Xiqi Hao, Zengqing Wu, Yu-Xuan Qiu +4
Beijing Institute of Technology, Zhuhai · University of Osaka · Shenzhen University
Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior work has mostly studied this in single-agent settings, it remains underexplored in collaborative multi-agent systems. We ask whether awareness of other agents' sycophancy levels influences discussion outcomes. To investigate this, we run controlled experiments with six open-source LLMs, providing agents with peer sycophancy rankings that estimate each peer's tendency toward sycophancy. These rankings are based on scores calculated using various static (pre-discussion) and dynamic (online) strategies. We find that providing sycophancy priors reduces the influence of sycophancy-prone peers, mitigates error-cascades, and improves final discussion accuracy by an absolute 10.5%. Thus, this is a lightweight and efficient way to reduce model sycophancy during discussions and subsequently improve downstream accuracy.