Autonomous agents powered by large language models (LLMs) are rapidly evolving into an open agentic ecosystem. To support trustworthy collaboration, industry initiatives increasingly assess agent reputation from past behavior and provide performance leaderboards. However, reputation derived from past performance may not reliably predict an agent's behavior on new tasks, particularly when malicious agents can adapt their behavior and influence other agents during collaboration. We study reputation in multi-agent debate (MAD), where multiple agents answer the same query, debate to improve their answers, and aggregate them into a final output. We present MiniRep, a reputation-based aggregation system for MAD under malicious agents. To ground our threat model in established research, we construct an attack taxonomy drawing on reputation-system attacks and software-testing mutation operators, covering strategic exploitation of reputation and subtle corruption of agent proposals. Guided by this taxonomy, MiniRep evaluates agents based on both their behavior on the current task and their reputation over time, while preventing groups of agents with highly similar responses from dominating the final decision. We assess MiniRep across diverse tasks, LLM-agent compositions, corruption placements, and attack types drawn from our taxonomy. Our experimental results show that, MiniRep outperforms both conventional MAD aggregation and conventional reputation-based approaches on MATH no matter being attacked or not. Also, under a heterogeneous 10-agent setting on MATH, MiniRep outperforms all baselines in all 28 attack conditions.
Figures & tables
Attack
Instantiation
Reputation class
Related analogue
Random coordinated
Malicious agents submit the same randomly selected incorrect answer.
Orchestrated
Adversarial peers can induce conformity and degrade agent decisions ( Ko et al., 2026 ) ; incorrect proposals can propagate during MAD ( Cui et al., 2026 ) .
Optimized coordinated
Malicious agents search for an incorrect proposal that maximizes its estimated impact on the final output.
Orchestrated
Adversarial agents can generate candidate arguments and select the most persuasive one ( Kraidia et al., 2026 ) .
Diverse collusion
Malicious agents submit distinct proposals that support the same incorrect conclusion.
Orchestrated
Colluding agents can provide distinct contributions toward a common objective and induce false conclusions ( Hu et al., 2026 ) .
On-off
Malicious agents first build reputation and then attack on every subsequent task.
Self-promoting ⋆
Agents can first build peer trust and later inject fabricated information ( Park et al., 2026 ) .
Adaptive
Malicious agents first build reputation. During an attack, only a selected subset deviates from the system specification.
Self-promoting ⋆
Manipulating one agent can be sufficient to mislead a multi-agent system ( Liu et al., 2025 ) .
Arithmetic operator
A malicious agent changes one arithmetic operation in an honestly generated proposal.
N/A
Arithmetic and relational operator replacement can introduce small but consequential faults ( Petrović and Ivanković, 2018 ) .
Table 1: Attack taxonomy studied in this work. Reputation classes follow Hoffman et al. (2009) . The last column lists related mechanisms reported in prior work; these studies do not explicitly target agent reputation. ⋆ On-off and Adaptive adapt self-promoting attacks to MAD.
Figure 1: Overview of MiniRep .
Name (#)
Settings
Dataset (3)
GoEmotions ( Demszky et al., 2020 ) , MATH (levels 4–5) ( Hendrycks et al., 2021 ) , HumanEval Pro ( Yu et al., 2025 )
Composition (4)
all strong, 7-strong + 3-weak (7S+3W), all weak, random
Placement (4)
B (top-ranked), W (bottom-ranked), R1, R2 (random)
Attacks (7 + clean)
Rand ( Hoffman et al., 2009 ; Ko et al., 2026 ) , Strong ( Kraidia et al., 2026 ) , Div ( Hu et al., 2026 ) , OnOff ( Park et al., 2026 ) , Adapt ( Hoffman et al., 2009 ; Liu et al., 2025 ) , Arith ( Petrović and Ivanković, 2018 ) , Bound ( Just et al., 2014 )
Table 2: Experimental setup. Attack abbreviations follow the order in Table 1 .
Clean
Attacked
Dataset / metric
Best baseline
MiniRep
Best baseline
MiniRep
GoEmotions / accuracy
21.50 (Beta)
21.75
19.59 (Beta)
21.16
GoEmotions / sample F1
36.42 (UMaj)
34.42
34.87 (Beta)
33.94
MATH / accuracy
59.25 (UMaj)
66.75
54.37 (Beta)
61.95
HumanEval Pro / Pass@1
79.50 (UMaj)
77.75
77.92 (Eigen)
78.15
Table 3: Average task scores (%). The best evaluated baseline is selected separately for results in the clean and attacked runs. Bold marks the better score; full eight-method results are in Table 6 .
Setting
Best baseline b
MiniRep
GoEmotions Acc./7S+3W
0/28 (all)
25/28
GoEmotions F1 / All weak
8/28 (Single)
17/28
MATH / All strong
0/28 (all)
27/28
MATH / 7S+3W
0/28 (all)
28/28
HumanEval Pro / 7S+3W
10/28 (Eigen)
12/28
Table 4: Performance summary against the best evaluated baseline b . Panel (a) reports strict wins out of 28 attack conditions for each composition. Panels (b)–(c) report attacked task scores (%) where bold numbers denote the better value. The complete data are included in Tables 8 – 10 , where repeated cells are highlighted.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Protocol related
P ; n ; f
Identity set; number of identities; maximum adversarial identities.
τ=(ctx,χ)
Task: context ctx (statement and public evidence) and parameter χ .
vi=(ci,ρi)
Response of identity i : payload ci and rationale ρi .
R
Debate rounds ( R=0 in the main setting).
G ; clone-group mapping
Known group of identities; the experiments group replicas by their underlying API model.
Appendix
Table 5: Core notation.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
Clean ↑
18.50
17.50
20.25
20.50
17.75
21.50
20.50
21.75
Attacked ↑
10.71
11.04
19.04
18.05
12.42
19.59
17.79
21.16
Δ↓
7.79
6.46
1.21
2.45
5.33
1.91
2.71
0.59
RelDrop ↓
42.64
37.30
6.41
12.93
30.33
9.25
13.54
2.52
GoEmotions / sample F1
Appendix
Table 6: Average performance for all eight methods. Clean scores are averaged over the four compositions. Attacked scores, paired losses, and relative drops are the average of 112 attack conditions. Scores and RelDrop are percentages. Δ is in percentage points. Shaded cells appear in Table 3 ; bold marks the best value in each row.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions
PoolExp ↓
100.00
95.41
53.72
46.35
95.49
47.43
53.20
81.79
NextExcl ↑
0.00
4.78
49.84
55.26
4.71
53.64
47.93
19.42
PayloadHit ↓
27.93
24.62
19.55
22.48
25.22
19.83
20.69
18.22
CF-Harm ↓
27.84
21.02
9.10
9.64
22.67
7.89
10.51
5.82
MATH
Appendix
Table 7: Average defense rates of all experiments per dataset.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
All strong
0/28
0/28
0/28
0/28
0/28
6/28
0/28
15/28
7 strong + 3 weak
0/28
0/28
0/28
0/28
0/28
0/28
0/28
25/28
All weak
0/28
0/28
2/28
0/28
0/28
2/28
0/28
18/28
Random
0/28
0/28
1/28
3/28
0/28
9/28
0/28
6/28
GoEmotions / sample F1
Appendix
Table 8: Strict wins for each agent composition. Bold marks the largest count in each row. Shaded cells are summarized in panel (a) of Table 4 .
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / B / Strong (F1; b= Beta; Rb=+0.4 )
Clean ↑
37.40
34.50
37.37
37.10
36.10
38.27
38.43
35.00
Attacked ↑
18.63
24.33
35.27
32.30
21.90
35.83
35.33
33.00
Δ↓
18.77
10.17
2.10
4.80
14.20
2.43
3.10
2.00
CF-Harm ↓
48.28
31.03
6.90
12.64
36.78
8.05
9.20
3.45
GoEmotions / B / Arith (F1; b= Beta; Rb=+4.2 )
Appendix
Table 9: Paired stress cases in the 7S+3W composition. Baseline b has the highest attacked score, where ties are resolved as mentioned in Appendix C.3.1 . Scores and CF-Harm are percentages; Δ and Rb are percentage points. Shaded cells are reported in Table 4 (b). Bold marks the best value in each row.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
HumanEval Pro / Pass@1 (shared clean reference)
Clean ↑
83.00
80.00
79.00
78.00
79.00
79.00
77.00
79.00
HumanEval Pro / B / OnOff ( b= Eigen)
Attacked ↑
45.00
51.00
77.00
75.00
78.00
75.00
51.00
81.00
CF-Harm ↓
54.29
41.43
5.71
7.14
4.29
8.57
40.00
1.43
PayloadHit ↓
64.29
52.94
16.67
44.44
14.29
35.71
54.29
0.00
Appendix
Table 10: Performance under OnOff and Adapt attacks on HumanEval Pro under the 7S+3W composition. Baseline b has the highest attacked Pass@1 in each setting, where ties are resolved as mentioned in Appendix C.3.1 . Shaded cells are reported in Table 4 (c). “–” denotes undefined PayloadHit. Bold marks the best defined value in each row.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
First five
11.59
11.59
19.55
18.54
12.91
19.76
19.00
21.33
Arith
9.00
9.50
18.44
17.31
12.31
20.00
15.06
20.94
Bound
8.00
9.81
17.06
16.38
10.06
18.31
14.44
20.56
GoEmotions / sample F1
First five
28.04
27.82
34.72
34.09
28.44
35.06
34.71
33.95
Appendix
Table 11: Test performance on the five attack families represented in feedback-model training and the two held-out families. First five averages 4×4×5=80 conditions; each held-out family averages 4×4=16 . Scores are percentages; bold marks the largest mean in each row. These are complete-system comparisons, not feedback-removal ablations.
Semantic low-quality
Behavior probe
Fused detector
Dataset / attack
P
N
Recall
FPR
Recall
FPR
GoEmotions
First five
91.40
60.03
23.55
9.69
24.77
12.30
Arith
78.51
59.39
21.73
9.58
25.18
12.92
Bound
80.61
57.99
21.68
9.74
26.34
13.05
MATH
Appendix
Table 12: Recorded proposal-level feedback diagnostics (%). Counts are pooled within each dataset and attack group before computing rates; each configuration is counted once. P and N denote proposals with and without an active attack payload. Probe and fused-detector recall/FPR use the active-payload label. The semantic columns measure low-quality predictions: N can include naturally incorrect proposals, so its low-quality rate is not a false-positive rate.
Metric / setting
UMaj
URand
Single
Babylon
Eigen
Beta
TS
MiniRep
GoEmotions / exact-match accuracy
All strong
13.14
14.18
21.46
21.46
14.64
22.71
21.21
23.79
7 strong + 3 weak
10.32
10.00
19.71
17.82
11.89
19.93
18.54
23.50
All weak
6.89
7.89
14.07
11.93
8.82
13.61
11.61
15.82
Random
12.46
12.07
20.89
21.00
14.32
22.11
19.79
21.54
GoEmotions / sample F1
Appendix
Table 13: Mean attacked task scores (%) under all four agent compositions. Every entry averages the four corruption placements and seven attacks ( 28 conditions), with the method configuration held fixed. Bold marks the largest mean in each row. Unlike the strict-win counts in Table 8 , this table retains score magnitudes.
School of Computer Science & Technology, Beijing Jiaotong University, Beijing, China · Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing, China