Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.
Figures & tables
Figure 1: Overview of our artifact-mediated propagation across independent assistants. A poisoned artifact infects an assistant’s memory ( hop 1 ). The assistant later reproduces the adversarial state into another artifact, which can infect a second assistant ( hop 2 ) and continue across further hops.
Target model
FIFagoal
FIFafull
FIFdgoal
S(1)
S(2)
S(3)
S(4)
GPT-5.6 Luna
0.38
0.32
0.33
0.80
0.57
0.40
0.24
Kimi-K2.6
0.47
0.37
0.50
0.78
0.67
0.61
0.44
GPT-OSS-120B
0.85
0.73
0.83
0.99
0.92
0.88
0.61
DeepSeek-V4-Pro
0.98
0.71
0.97
1.00
0.93
0.93
0.76
Table 1: Main propagation results on the held-out test universes. Results are averaged over repeated runs within each universe and then across universes.
Figure 2: Goal-infection hop survival on the held-out test universes. Dotted lines show observed S(h) ; solid lines show the censoring-corrected Kaplan–Meier estimate with 95% bootstrap confidence intervals (Section 5.3 ). Dashed lines show a geometric extrapolation using the fitted per-hop continuation probability p^ ; each solid curve ends at the deepest observed hop.
Figure 3: Decomposing a single exposure over verified-clean reads of judged artifacts. Encounter is the probability that a clean read lands on an artifact carrying the adversarial state; conversion is the probability that the reader then acquires the goal. Whiskers: 95% CIs.
Figure 4: Contact networks of the 12 large-universe runs. Rows are universes and columns are target models; within a row, each agent occupies the same position in every panel, so the panels differ only in how the models behave. Filled nodes were infected at some point during the run and hollow nodes never were. Stars mark the three seed agents, hollow if the seed agent never converted. Thin grey lines connect agents that read each other’s artifacts, and colored lines are the attributed transmissions from parent to child. GPT-5.6 Luna leaves a clean periphery in every universe, while GPT-OSS-120B and DeepSeek-V4-Pro reach nearly every agent from the same starting points.
Figure 5: (a) Number of agents each infected agent infects directly, pooled over the 12 large-universe runs. (b) Share of all attributed transmissions caused by the most prolific infected agents, by model; the dashed line is the curve if every infected agent contributed equally, and the grey vertical line marks the top 20%.
Figure 6: Transmission tree rooted at a single artifact in the DeepSeek-V4-Pro run on large universe A. Each node is an infected agent placed at its time of infection; dashed lines are direct reads of the artifact, and solid lines are onward transmissions to other agents. The artifact, written by seed agent 23 at step 5, directly infects 14 agents over the following 41 steps.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Level
Label
Active fraction
Mean degree
Hand-offs / agent
Group size
Length
1
Light
0.22–0.33
2.2
1.2
3–10 / 6.0
10–10 / 10.0
2
Steady
0.17–0.33
2.1
1.3
3–10 / 6.0
11–13 / 12.2
3
Busy
0.19–0.26
3.5
2.2
5–10 / 7.5
14–16 / 14.3
4
Dense
0.22–0.28
3.5
2.3
5–11 / 8.2
12–15 / 14.5
5
Very dense
0.16–0.28
3.6
2.2
6–11 / 8.5
13–16 / 14.7
6
Peak
0.36–0.40
4.1
4.6
7–12 / 9.5
20–20 / 20.0
Appendix
Table 2: The six workload-intensity levels (6 universes each). Active fraction is a range over universes; mean degree and hand-offs per agent are means over universes; group size and length are range / mean.
Domain
S(1)
S(2)
S(3)
Healthcare-operations
0.60
0.50
0.20
Customer-support
0.50
0.40
0.10
Software-engineering
0.50
0.30
0.10
Home-living
0.30
0.30
0.10
Social-gatherings
0.20
0.20
0.10
Personal-productivity
0.20
0.10
0.10
Appendix
Table 3: Held-out prompt-only propagation on DeepSeek-V4-Flash , by domain (full infection; 30 universes, two runs each; observed S(h) ).
Hop
Artifacts
Goal verbatim
Instruction verbatim
Drop rate
1
50
80%
72%
32%
2
14
64%
79%
65%
3
5
60%
80%
64%
Appendix
Table 4: Held-out per-hop faithfulness among payload-carrying artifacts. The drop rate is over artifacts written by an agent already holding the payload in memory.
Target model
FIFagoal
FIFafull
FIFdgoal
S(1)
S(2)
S(3)
S(4)
Main results: goals sampled across all categories (Table 1 )
GPT-5.6 Luna
0.38
0.32
0.33
0.80
0.57
0.40
0.24
Kimi-K2.6
0.47
0.37
0.50
0.78
0.67
0.61
0.44
GPT-OSS-120B
0.85
0.73
0.83
0.99
0.92
0.88
0.61
DeepSeek-V4-Pro
0.98
0.71
0.97
1.00
0.93
0.93
0.76
Security-critical goals only ( action category)
Appendix
Table 5: Propagation of security-critical ( action -category) adversarial goals compared with the category-mixed main results, on the same 36 held-out test universes and the same frozen, goal-agnostic templates. The security-critical runs use one run per universe rather than two.
Target model
Jobs
Target cost (USD)
Outcome
GPT-OSS-120B
109
1.71
stopping rule met
Kimi-K2.6
372
33.18
stopping rule met
DeepSeek-V4-Pro
932
89.44
stopping rule met
GPT-5.6 Luna
4,002
198.58
stopping rule met
Grok 4.6
571
58.97
stopped for cost
Appendix
Table 6: Development effort per target model. Jobs are runs of a candidate template on one development universe; cost counts only calls to the target model. Grok 4.6 was stopped before meeting the stopping rule.
Figure 7: Second-hop survival S(2) of candidate templates over the course of the search for the weaker models and DeepSeek-V4-Pro. Bottom axis: cumulative jobs; top axis: cumulative target-model cost, with ticks placed at the job count where each cost was reached, so their spacing is uneven. Points are individual candidate templates and the black line is a rolling average. The dashed line marks the S(2)>0.50 threshold of the stopping rule (Section 4.2 ).
Figure 8: Second-hop survival S(2) of candidate templates over the course of the search for GPT-5.6 Luna and the frontier-class Grok 4.6. Bottom axis: cumulative jobs; top axis: cumulative target-model cost, with ticks placed at the job count where each cost was reached, so their spacing is uneven. Points are individual candidate templates and the black line is a rolling average. The dashed line marks the S(2)>0.50 threshold of the stopping rule (Section 4.2 ).
Figure 9: Fraction of agents goal-infected at the end of a run, by workload intensity (a) and by domain (b), for each target model.
Figure 10: Cumulative fraction of agents that have been goal-infected at each time step in the 30-agent universes, averaged over the three universes for each target model. The dotted line marks half of the population. Full infection differs from goal infection by at most two agents in any run, and not at all in 10 of the 12 runs, so only goal infection is shown.
Figure 11: (a) Composition of each run’s 30 agents: the largest single seed lineage, the lineages of the other two seeds, infected agents whose source could not be attributed, and agents that were never infected. In most runs one lineage accounts for the majority of infections. (b) Never-infected agents by model, split into those that read at least one judged-infected artifact and resisted (solid) and those that never read one (hatched).
Figure 12: Cumulative infections in each seed artifact’s lineage for the GPT-OSS-120B run on universe C.
Figure 13: Transmission tree for DeepSeek-V4-Pro on universe B, in which two of the three seed agents never convert. Each node is an infected agent, placed at its time of infection and connected to its attributed parent. Seed agent 06 reads its seed artifact and never acquires the goal, but the artifact is later read directly by agents 12 and 14 and by seed agent 13 (dashed lines), each exposed exactly as a seed agent would be. Their lineages account for 25 of the run’s 28 infections. The three agents below the line have no attributed source.
Figure 14: Lineage of seed agent 23 in the Kimi-K2.6 run on universe A, with the longest attributed chain highlighted and numbered by hop. Each node is an infected agent placed at its time of infection.
Reproduction number R
Target model
Steps 0–15 ( n )
Steps 16–30
Dispersion k
Generation interval
GPT-5.6 Luna
1.32 (19)
0.94
>103
9
Kimi-K2.6
2.12 (25)
0.41
0.37
10
GPT-OSS-120B
1.57 (30)
0.57
1.18
10
DeepSeek-V4-Pro
1.59 (29)
0.68
0.62
10
All models
1.66
—
0.79
10
Appendix
Table 7: Epidemiological metrics for the 12 large-universe runs, computed from the attributed transmission trees and pooled over the three universes of each model. n is the number of agents infected by time step 15.
Figure 15: Epidemiological metrics for the large-universe runs. (a) Reproduction number for agents infected in the first 15 time steps and in steps 16–30, with 95% bootstrap confidence intervals; the dashed line marks R=1 . (b) Fraction of infected agents that infect 0, 1, 2, 3, or 4 or more others, with the fitted dispersion k in the legend. (c) Generation intervals, the time from an agent’s infection to each infection it causes.
Test universes
Large universes
Target model
Upload rate
Docs/universe
Upload rate
Docs/universe
GPT-5.6 Luna
0.96
4.2
1.00
29.0
Kimi-K2.6
0.97
7.0
0.96
42.7
GPT-OSS-120B
0.97
13.7
0.98
73.0
DeepSeek-V4-Pro
1.00
15.8
1.00
63.0
Appendix
Table 8: Exfiltration under the endpoint-assisted mechanism. “Upload rate” is the fraction of infected agents that send at least one artifact to the endpoint; “Docs/universe” is the mean number of artifacts sent per universe.
Figure 16: How often an infected agent’s new artifact reproduces the belief, the propagation instruction verbatim, and the instruction including paraphrases, pooled over all writes on the test universes.
Memory lists
n
Flagged
Score range
Infected (should be flagged)
Main template, judge-verified
104
104
≥0.98
Command and URL lines removed
5
5
0.97–0.98
Held-out template variant
5
5
≥0.98
Instructions paraphrased by an LLM
12
12
0.96–0.99
Twelve distinct template styles
12
11
0.31–0.94
Appendix
Table 9: Classifier results with the generic question at threshold 0.5 . Flagged counts memory lists with score ≥0.5 ; the desired outcome is to flag every list in the upper block and none in the lower block.
Study
State carrying the attack
Continuation path
Goal evaluation
Hidden in Memory ( Pulipaka et al., 2026 )
Private agent memory
The same assistant uses the poisoned memory in later conversations; transfer to another assistant is not evaluated.
Reusable template across many goals.
MURMUR ( Patlan et al., 2025 )
Persistent state shared by users of one collaborative agent
A later user interacts with that same agent and is affected by its shared state.
Attacker-specified actions.
Prompt Infection ( Lee & Tiwari, 2024 ; Yu et al., 2024 )
Agent messages and active context
Infected agents pass instructions directly to other agents in a connected multi-agent system.
Attack-specific payloads.
AI Worm ( Cohen et al., 2025 )
Email indexed by an assistant’s RAG system
An assistant reproduces the prompt in outgoing email, which can enter another assistant’s email and RAG pipeline.
Predefined malicious effects.
AgentWorm ( Zhang et al., 2026b )
Agent configuration and other persistent files
An infected agent autonomously reaches peers through messaging and agent ecosystem channels.
Three payload types.
Autonomous LLM Agent Worms ( Zha & Wang, 2026 )
File-backed workspace and memory state
Persistent content re-enters execution and drives autonomous cross-platform transmission.
Selected malicious actions.
Appendix
Table 10: Attack paths in closely related work. The final column describes the evaluated payload or template; it does not equate transfer across tasks with transfer across adversarial goals.
‡∗§Department of Computer Science, §Klipsch School of Electrical and Computer Engineering New Mexico State University, Las Cruces, New Mexico, USA · §Klipsch School of Electrical and Computer Engineering New Mexico State University, Las Cruces, New Mexico, USA