Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agents, whose semantic states are dynamically refined through a recurrent process known as collaborative reflection. While this mechanism improves recommendation quality, it simultaneously introduces a systemic vulnerability: adversarial evidence injected into a single agent can be rationalised into a legitimate preference narrative, written back into memory, and propagated to other agents through interaction contexts. We term the local rationalisation process reflection laundering, and its system-wide escalation through collaborative reflection collaborative-reflection hijacking. Existing attacks on recommender systems, whether based on interaction-level data poisoning or text-level adversarial perturbations, assume static pipelines and thus cannot exploit this recurrent, multi-agent amplification pathway. To bridge this gap, we first conduct a controlled vulnerability analysis that establishes two exploitable properties underlying collaborative-reflection hijacking: reflective persistence and cross-agent propagation. Then building on these findings, we propose VirusCascade, the first black-box targeted promotion attack that jointly shapes semantic and structural attack surfaces: the former ensures the target item is naturally rationalised as satisfying broad user preferences, the latter positions it for system-wide propagation. Extensive experiments on four real-world datasets across diverse LLM-ARS architectures demonstrate that VirusCascade consistently achieves state-of-the-art targeted exposure under evaluated stealth constraints, reaching a mean E@20 of 0.384 and surpassing the strongest baseline by an absolute margin of +0.185.
Figures & tables
Fig. 1: RS vs. LLM-ARS. Traditional RS (Left) optimises fixed user–item representations through gradient-based learning, whereas LLM-ARS instantiate users and items as memory-equipped agents that update their semantic states through autonomous interaction and collaborative reflection.
TABLE I: Reflective persistence among admitted users (AgentCF, CDs & Vinyl), averaged over 10 runs. Each column reports a lag relative to the user’s individual admission round.
Fig. 2: Cumulative fraction of benign users measured across interaction-reflection rounds. Dark bars: newly admitted in each round; light bars: admitted in prior rounds and still retained. Dashed red line: no-probe baseline regeneration rate (2%).
Users
n
Setting
Rec. Drift ↑
Claim Adoption ↑
Hop-1
12
Sem.-null
0.003 [0.001, 0.005]
0.0% [0.0, 0.0]
Sem.-probe
0.397 [0.371, 0.423]
33.6% [28.9, 38.3]
Hop-2
25
Sem.-null
0.002 [0.001, 0.003]
0.0% [0.0, 0.0]
Sem.-probe
0.245 [0.221, 0.269]
25.8% [22.7, 28.9]
Hop >2
63
Sem.-null
0.003 [0.002, 0.004]
0.0% [0.0, 0.0]
Sem.-probe
0.123 [0.112, 0.134]
6.7% [5.1, 8.3]
TABLE II: Cross-agent propagation under the semantic probe and topology-matched semantic-null control. Values report the mean and 95% t -confidence interval across 10 paired runs; n denotes the number of benign users in each hop group.
Fig. 3: Overview of VirusCascade . The adversary extracts preference motifs from anchor items and designs the target profile to make i∗ easier to rationalise during reflection. It then constructs smooth, propagation-oriented behaviour trajectories from anchor items to i∗ . The victim system’s collaborative reflection may subsequently admit and propagate the injected evidence.
Dataset
#Users
#Items
#Inter.
Sparsity
CDs & Vinyl
112,395
73,713
1,443,755
99.98%
Movies & TV
297,529
60,110
3,404,812
99.98%
Automotive
193,651
79,317
1,709,025
99.99%
Musical Instruments
27,530
10,611
231,312
99.92%
TABLE III: Statistics of the datasets used in our experiments.
Victim RS
AgentCF
Attack
CDs & Vinyl
Movies & TV
H@10
N@10
E@10
H@20
N@20
E@20
H@10
N@10
E@10
H@20
N@20
E@20
NoAttack
0.130
0.050
(-)
0.260
0.082
(-)
0.220
0.125
(-)
0.310
0.148
(-)
RandAttack
0.090
0.053
0.010
0.210
0.084
0.071
0.170
0.110
0.050
0.280
0.138
0.120
PopAttack
0.115
0.053
0.000
0.235
0.077
0.010
0.170
0.114
0.050
0.290
0.133
0.150
ExpPromot
0.140
0.054
0.000
0.290
0.091
0.051
0.185
0.117
0.030
0.270
0.134
0.110
TABLE IV: Attack performance across three victim architectures on CDs & Vinyl and Movies & TV datasets. We report H@K and N@K for recommendation utility and E@K for targeted promotion effectiveness. NoAttack denotes the same victim configuration and benign protocol without any adversarial manipulation, and we report its E@K as (-).
Dataset
NoAttack
DrunkAgent
Ours
Range
Mean
CDs & Vinyl
[10.31, 90.02]
28.78
49.93
32.22
Movies & TV
[12.91, 374.28]
72.85
50.65
39.35
TABLE V: Attack imperceptibility in terms of perplexity.
Attacks
CDs & Vinyl
Movies & TV
@1
@2
@L
@1
@2
@L
TextFooler
0.200
0.000
0.200
0.251
0.000
0.121
DrunkAgent
0.294
0.000
0.100
0.175
0.000
0.087
VirusCascade
0.321
0.109
0.268
0.294
0.090
0.157
TABLE VI: Attack imperceptibility in terms of ROUGE.
Attacks
Flagged as Manipulated ↓
Factual Consistency ↑
Original
19.3%
4.31±0.47
DrunkAgent
28.7%
4.02±0.95
VirusCascade
21.9%
4.16±0.62
TABLE VII: Attack imperceptibility in terms of human evaluation (20 participants, 120 descriptions each).
Victim RS
AgentCF
Variant
CDs & Vinyl
Movies & TV
H@10
N@10
E@10
H@20
N@20
E@20
H@10
N@10
E@10
H@20
N@20
E@20
NoAttack
0.130
0.050
(-)
0.260
0.082
(-)
0.220
0.125
(-)
0.310
0.148
(-)
w.o SemInj
0.110
0.058
0.010
0.250
0.094
0.071
0.200
0.135
0.020
0.320
0.164
0.170
w.o StruInj
0.100
0.046
0.020
0.260
0.086
0.131
0.160
0.101
0.000
0.300
0.135
0.080
VirusCascade
0.120
0.057
0.111
0.250
0.083
0.374
0.205
0.123
0.070
0.305
0.138
0.240
TABLE VIII: Ablation study: effect of removing semantic injection (w.o SemInj) or structural injection (w.o StruInj).
Fig. 4: Effect of malicious user fraction m .
Fig. 5: Effect of the coefficient α .
Dataset
Victim RS
LLaMA-3
GPT-4o
Gemini-2.5
H@20
N@20
E@20
H@20
N@20
E@20
H@20
N@20
E@20
CDs & Vinyl
AgentCF
0.320
0.107
0.151
0.250
0.083
0.374
0.340
0.119
0.333
AgentSEQ
0.410
0.129
0.160
0.240
0.087
0.465
0.330
0.112
0.283
AgentRAG
0.320
0.096
0.170
0.230
0.105
0.545
0.350
0.115
0.364
Movies & TV
AgentCF
0.250
0.102
0.140
0.290
0.134
0.240
0.340
0.161
0.140
AgentSEQ
0.330
0.144
0.150
0.300
0.142
0.300
0.340
0.165
0.210
TABLE IX: Attack transferability with different auxiliary models Maux .
Dataset
Maux
PPL
ROUGE
NoAttack
Ours
@1
@2
@L
CDs & Vinyl
LLaMA-3
[21.64, 107.62]
56.28
0.267
0.103
0.267
GPT-4o
[10.31, 90.02]
32.22
0.321
0.109
0.268
Gemini-2.5
[11.78, 113.24]
31.82
0.341
0.125
0.268
Movies & TV
LLaMA-3
[13.75, 409.59]
76.20
0.412
0.091
0.206
GPT-4o
[12.91, 374.28]
39.35
0.294
0.090
0.157
TABLE X: Attack imperceptibility with different auxiliary models Maux . NoAttack PPL is reported as [min, max].
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Marker
No-Probe Regen.
Lag-1
Lag-2
Lag-3
Lag-4
Rare
2.0%
78.0%
81.7%
79.3%
81.6%
Common
15.0%
72.0%
73.4%
68.7%
76.4%
Appendix
TABLE XI: Reflective persistence under rare and common-frequency diagnostic markers. Lag columns report reuse rate.
Victim RS
AgentCF (user-only)
Attack
CDs & Vinyl
Movies & TV
H@10
N@10
E@10
H@20
N@20
E@20
H@10
N@10
E@10
H@20
N@20
E@20
NoAttack
0.170
0.074
(-)
0.350
0.119
(-)
0.190
0.124
(-)
0.320
0.157
(-)
TextBugger
0.120
0.057
0.010
0.370
0.119
0.061
0.220
0.133
0.020
0.320
0.158
0.130
TextFooler
0.150
0.063
0.040
0.340
0.111
0.051
0.210
0.115
0.010
0.330
0.145
0.110
DrunkAgent
0.140
0.069
0.000
0.270
0.102
0.030
0.220
0.129
0.020
0.340
0.160
0.100
Appendix
TABLE XII: Attack transferability under user-only agent architectures (no item-side memory updates).
Victim RS
AgentCF
Attack
Automotive
Musical Instruments
H@10
N@10
E@10
H@20
N@20
E@20
H@10
N@10
E@10
H@20
N@20
E@20
NoAttack
0.150
0.075
(-)
0.230
0.094
(-)
0.110
0.042
(-)
0.280
0.085
(-)
RandAttack
0.080
0.043
0.000
0.180
0.068
0.170
0.060
0.023
0.010
0.160
0.049
0.141
PopAttack
0.070
0.038
0.050
0.230
0.078
0.170
0.090
0.043
0.030
0.220
0.075
0.091
ExpPromot
0.100
0.043
0.030
0.230
0.076
0.180
0.030
0.011
0.051
0.150
0.040
0.152
Appendix
TABLE XIII: Attack performance on Automotive and Musical Instruments. H@K and N@K measure recommendation utility, while E@K measures targeted promotion effectiveness.
Fig. 6: Population-scale evaluation on CDs & Vinyl under a fixed attacker budget of five users. Bars report VirusCascade as the active population increases from 100 to 500 users. Star markers denote DrunkAgent under the corresponding population setting.
Fig. 7: Attack performance over an extended 20-round interaction on AgentCF with CDs & Vinyl. Top: recommendation utility (H@20); bottom: targeted exposure (E@20).
Defence
Attack
AgentCF
AgentSEQ
AgentRAG
H@20
N@20
E@20
H@20
N@20
E@20
H@20
N@20
E@20
ONION (T)
NoAttack
0.410
0.141
(-)
0.410
0.161
(-)
0.410
0.168
(-)
DrunkAgent
0.310
0.107
0.030
0.380
0.135
0.030
0.400
0.131
0.040
VirusCascade
0.340
0.128
0.256
0.370
0.130
0.202
0.350
0.140
0.202
Fraudar (B)
NoAttack
0.270
0.088
(-)
0.380
0.124
(-)
0.340
0.114
(-)
DrunkAgent
0.290
0.109
0.061
0.390
0.145
0.081
0.380
0.138
0.061
Appendix
TABLE XIV: Attack performance under text- (T), behaviour- (B), and memory-level (M) defences on the CDs & Vinyl dataset. Comparisons are made within each defence block, as defences alter the inference pipeline.