Autonomous agents powered by large language models (LLMs) are rapidly evolving into an open agentic ecosystem. To support trustworthy collaboration, industry initiatives increasingly assess agent reputation from past behavior and provide performance leaderboards. However, reputation derived from past performance may not reliably predict an agent's behavior on new tasks, particularly when malicious agents can adapt their behavior and influence other agents during collaboration. We study reputation in multi-agent debate (MAD), where multiple agents answer the same query, debate to improve their answers, and aggregate them into a final output. We present MiniRep, a reputation-based aggregation system for MAD under malicious agents. To ground our threat model in established research, we construct an attack taxonomy drawing on reputation-system attacks and software-testing mutation operators, covering strategic exploitation of reputation and subtle corruption of agent proposals. Guided by this taxonomy, MiniRep evaluates agents based on both their behavior on the current task and their reputation over time, while preventing groups of agents with highly similar responses from dominating the final decision. We assess MiniRep across diverse tasks, LLM-agent compositions, corruption placements, and attack types drawn from our taxonomy. Our experimental results show that, MiniRep outperforms both conventional MAD aggregation and conventional reputation-based approaches on MATH no matter being attacked or not. Also, under a heterogeneous 10-agent setting on MATH, MiniRep outperforms all baselines in all 28 attack conditions.
Figures & tables
Attack
Instantiation
Reputation class
Related analogue
Random coordinated
Malicious agents submit the same randomly selected incorrect answer.
Orchestrated
Adversarial peers can induce conformity and degrade agent decisions ( Ko et al., 2026 ) ; incorrect proposals can propagate during MAD ( Cui et al., 2026 ) .
Optimized coordinated
Malicious agents search for an incorrect proposal that maximizes its estimated impact on the final output.
Orchestrated
Adversarial agents can generate candidate arguments and select the most persuasive one ( Kraidia et al., 2026 ) .
Diverse collusion
Malicious agents submit distinct proposals that support the same incorrect conclusion.
Orchestrated
Colluding agents can provide distinct contributions toward a common objective and induce false conclusions ( Hu et al., 2026 ) .
On-off
Malicious agents first build reputation and then attack on every subsequent task.
Self-promoting ⋆
Agents can first build peer trust and later inject fabricated information ( Park et al., 2026 ) .
Adaptive
Malicious agents first build reputation. During an attack, only a selected subset deviates from the system specification.
Self-promoting ⋆
Manipulating one agent can be sufficient to mislead a multi-agent system ( Liu et al., 2025 ) .
Arithmetic operator
A malicious agent changes one arithmetic operation in an honestly generated proposal.
N/A
Arithmetic and relational operator replacement can introduce small but consequential faults ( Petrović and Ivanković, 2018 ) .
Table 1: Attack taxonomy studied in this work. Reputation classes follow Hoffman et al. (2009) . The last column lists related mechanisms reported in prior work; these studies do not explicitly target agent reputation. ⋆ On-off and Adaptive adapt self-promoting attacks to MAD.
Figure 1: Overview of MiniRep .
Name (#)
Settings
Dataset (3)
GoEmotions ( Demszky et al., 2020 ) , MATH (levels 4–5) ( Hendrycks et al., 2021 ) , HumanEval Pro ( Yu et al., 2025 )
Composition (4)
all strong, 7-strong + 3-weak (7S+3W), all weak, random
Placement (4)
B (top-ranked), W (bottom-ranked), R1, R2 (random)
Attacks (7 + clean)
Rand ( Hoffman et al., 2009 ; Ko et al., 2026 ) , Strong ( Kraidia et al., 2026 ) , Div ( Hu et al., 2026 ) , OnOff ( Park et al., 2026 ) , Adapt ( Hoffman et al., 2009 ; Liu et al., 2025 ) , Arith ( Petrović and Ivanković, 2018 ) , Bound ( Just et al., 2014 )
Table 2: Experimental setup. Attack abbreviations follow the order in Table 1 .
Clean
Attacked
Dataset / metric
Best baseline
MiniRep
Best baseline
MiniRep
GoEmotions / accuracy
21.50 (Beta)
21.75
19.59 (Beta)
21.16
GoEmotions / sample F1
36.42 (UMaj)
34.42
34.87 (Beta)
33.94
MATH / accuracy
59.25 (UMaj)
66.75
54.37 (Beta)
61.95
HumanEval Pro / Pass@1
79.50 (UMaj)
77.75
77.92 (Eigen)
78.15
Table 3: Average task scores (%). The best evaluated baseline is selected separately for results in the clean and attacked runs. Bold marks the better score; full eight-method results are in Table 6 .
Setting
Best baseline b
MiniRep
GoEmotions Acc./7S+3W
0/28 (all)
25/28
GoEmotions F1 / All weak
8/28 (Single)
17/28
MATH / All strong
0/28 (all)
27/28
MATH / 7S+3W
0/28 (all)
28/28
HumanEval Pro / 7S+3W
10/28 (Eigen)
12/28
Table 4: Performance summary against the best evaluated baseline b . Panel (a) reports strict wins out of 28 attack conditions for each composition. Panels (b)–(c) report attacked task scores (%) where bold numbers denote the better value. The complete data are included in Tables 8 – 10 , where repeated cells are highlighted.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Protocol related
P ; n ; f
Identity set; number of identities; maximum adversarial identities.
τ=(ctx,χ)
Task: context ctx (statement and public evidence) and parameter χ .
vi=(ci,ρi)
Response of identity i : payload ci and rationale ρi .
R
Debate rounds ( R=0 in the main setting).
G ; clone-group mapping
Known group of identities; the experiments group replicas by their underlying API model.
Appendix
Table 5: Core notation.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
Clean ↑
18.50
17.50
20.25
20.50
17.75
21.50
20.50
21.75
Attacked ↑
10.71
11.04
19.04
18.05
12.42
19.59
17.79
21.16
Δ↓
7.79
6.46
1.21
2.45
5.33
1.91
2.71
0.59
RelDrop ↓
42.64
37.30
6.41
12.93
30.33
9.25
13.54
2.52
GoEmotions / sample F1
Appendix
Table 6: Average performance for all eight methods. Clean scores are averaged over the four compositions. Attacked scores, paired losses, and relative drops are the average of 112 attack conditions. Scores and RelDrop are percentages. Δ is in percentage points. Shaded cells appear in Table 3 ; bold marks the best value in each row.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions
PoolExp ↓
100.00
95.41
53.72
46.35
95.49
47.43
53.20
81.79
NextExcl ↑
0.00
4.78
49.84
55.26
4.71
53.64
47.93
19.42
PayloadHit ↓
27.93
24.62
19.55
22.48
25.22
19.83
20.69
18.22
CF-Harm ↓
27.84
21.02
9.10
9.64
22.67
7.89
10.51
5.82
MATH
Appendix
Table 7: Average defense rates of all experiments per dataset.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
All strong
0/28
0/28
0/28
0/28
0/28
6/28
0/28
15/28
7 strong + 3 weak
0/28
0/28
0/28
0/28
0/28
0/28
0/28
25/28
All weak
0/28
0/28
2/28
0/28
0/28
2/28
0/28
18/28
Random
0/28
0/28
1/28
3/28
0/28
9/28
0/28
6/28
GoEmotions / sample F1
Appendix
Table 8: Strict wins for each agent composition. Bold marks the largest count in each row. Shaded cells are summarized in panel (a) of Table 4 .
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / B / Strong (F1; b= Beta; Rb=+0.4 )
Clean ↑
37.40
34.50
37.37
37.10
36.10
38.27
38.43
35.00
Attacked ↑
18.63
24.33
35.27
32.30
21.90
35.83
35.33
33.00
Δ↓
18.77
10.17
2.10
4.80
14.20
2.43
3.10
2.00
CF-Harm ↓
48.28
31.03
6.90
12.64
36.78
8.05
9.20
3.45
GoEmotions / B / Arith (F1; b= Beta; Rb=+4.2 )
Appendix
Table 9: Paired stress cases in the 7S+3W composition. Baseline b has the highest attacked score, where ties are resolved as mentioned in Appendix C.3.1 . Scores and CF-Harm are percentages; Δ and Rb are percentage points. Shaded cells are reported in Table 4 (b). Bold marks the best value in each row.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
HumanEval Pro / Pass@1 (shared clean reference)
Clean ↑
83.00
80.00
79.00
78.00
79.00
79.00
77.00
79.00
HumanEval Pro / B / OnOff ( b= Eigen)
Attacked ↑
45.00
51.00
77.00
75.00
78.00
75.00
51.00
81.00
CF-Harm ↓
54.29
41.43
5.71
7.14
4.29
8.57
40.00
1.43
PayloadHit ↓
64.29
52.94
16.67
44.44
14.29
35.71
54.29
0.00
Appendix
Table 10: Performance under OnOff and Adapt attacks on HumanEval Pro under the 7S+3W composition. Baseline b has the highest attacked Pass@1 in each setting, where ties are resolved as mentioned in Appendix C.3.1 . Shaded cells are reported in Table 4 (c). “–” denotes undefined PayloadHit. Bold marks the best defined value in each row.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
First five
11.59
11.59
19.55
18.54
12.91
19.76
19.00
21.33
Arith
9.00
9.50
18.44
17.31
12.31
20.00
15.06
20.94
Bound
8.00
9.81
17.06
16.38
10.06
18.31
14.44
20.56
GoEmotions / sample F1
First five
28.04
27.82
34.72
34.09
28.44
35.06
34.71
33.95
Appendix
Table 11: Test performance on the five attack families represented in feedback-model training and the two held-out families. First five averages 4×4×5=80 conditions; each held-out family averages 4×4=16 . Scores are percentages; bold marks the largest mean in each row. These are complete-system comparisons, not feedback-removal ablations.
Semantic low-quality
Behavior probe
Fused detector
Dataset / attack
P
N
Recall
FPR
Recall
FPR
GoEmotions
First five
91.40
60.03
23.55
9.69
24.77
12.30
Arith
78.51
59.39
21.73
9.58
25.18
12.92
Bound
80.61
57.99
21.68
9.74
26.34
13.05
MATH
Appendix
Table 12: Recorded proposal-level feedback diagnostics (%). Counts are pooled within each dataset and attack group before computing rates; each configuration is counted once. P and N denote proposals with and without an active attack payload. Probe and fused-detector recall/FPR use the active-payload label. The semantic columns measure low-quality predictions: N can include naturally incorrect proposals, so its low-quality rate is not a false-positive rate.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
All strong
13.14
14.18
21.46
21.46
14.64
22.71
21.21
23.79
7 strong + 3 weak
10.32
10.00
19.71
17.82
11.89
19.93
18.54
23.50
All weak
6.89
7.89
14.07
11.93
8.82
13.61
11.61
15.82
Random
12.46
12.07
20.89
21.00
14.32
22.11
19.79
21.54
GoEmotions / sample F1
Appendix
Table 13: Mean attacked task scores (%) under all four agent compositions. Every entry averages the four corruption placements and seven attacks ( 28 conditions), with the method configuration held fixed. Bold marks the largest mean in each row. Unlike the strict-win counts in Table 8 , this table retains score magnitudes.
Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although some efforts have been made to examine particular attack types on MAD, systematic evaluation of MAD under diverse attacks remains limited. A central question is whether debate mitigates adversarial influence or amplifies it. In this paper, we present MADBench, a benchmark for evaluating the security of MAD. We organize attacks into a layered taxonomy following the MAD workflow, incorporating both established attacks and new strategies tailored to debate. We evaluate six attack families over 356 source tasks and 3,958 test cases, examining their effects on the final answer and the propagation of adversarial influence. Our results show that, under attacks, MAD does not necessarily improve LLM reasoning. Compared with a single-agent baseline, MAD can mitigate attacks on answer accuracy in question-answering tasks while amplifying unauthorized reads or writes in both question-answering and workspace tasks. Moreover, even when three out of five agents collude, the attack changes the final answer from correct to wrong on only 28.30% of tasks answered correctly without attack, while only 3.26% of initially correct honest agents switch to wrong answers during debate.
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the agents' inherently biased concept priors unaddressed. To mitigate this systematic weakness, we propose R2-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates. R2-MAD intervenes on both failure modes through two complementary mechanisms: A debate-state-aware retrieval policy dynamically calibrates the concept prior by retrieving relevant historical evidence based on the current consensus level. Then these retrieved experiences provide a basis for estimating per-agent reliability, yielding confidence weights to modulate peer influence. Experiments on various benchmarks show that R2-MAD achieves consistent improvements over existing single-agent and MAD baselines.
Xuanfa Jin, Zhijian Ma, Yongcheng Zeng +3
Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University College London
Multi-Agent Debate (MAD) improves the reasoning performance of Large Language Models (LLMs) through multi-round interaction. However, LLMs in MAD are highly susceptible to blind conformity. Existing individual evaluation methods, typically based on confidence or perplexity, fail to reflect the correctness of reasoning and may even exacerbate blind conformity. To address this, we shift the perspective from individual evaluation to group interaction. We define mutual referencing among LLMs as \textbf{Debate Relationships} and recognize that regulating these relationships is the key to mitigating blind conformity. In this paper, we propose a novel framework for \textbf{D}ynamically r\textbf{E}gulating deb\textbf{A}te \textbf{R}elationships (DEAR) from the group perspective. At first, DEAR quantifies consensus and divergence as \textit{group evidence} to capture the debate state. Then, DEAR operates through three stages: 1) What: perceiving group consultation tendency and uncertainty; 2) Who: introducing a Selection RL-Agent to dynamically select reference peers; and 3) How: adopting a Behavior RL-Agent to adaptively adjust generation behaviors. Notably, we formulate the execution of the two RL-Agents as a sequential decision-making process, jointly optimizing via multi-agent reinforcement learning. Extensive experiments demonstrate that DEAR achieves superior performance while significantly reducing token consumption.
Hao Wu, Shoucheng Song, Chang Yao +4
School of Computer Science & Technology, Beijing Jiaotong University, Beijing, China · Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing, China