Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents
Organizations: SPAR · University of Cambridge · APTA AI · CISPA Helmholtz Center for Information Security
Abstract
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.
Figures & tables
| Target model | |||||||
|---|---|---|---|---|---|---|---|
| GPT-5.6 Luna | 0.38 | 0.32 | 0.33 | 0.80 | 0.57 | 0.40 | 0.24 |
| Kimi-K2.6 | 0.47 | 0.37 | 0.50 | 0.78 | 0.67 | 0.61 | 0.44 |
| GPT-OSS-120B | 0.85 | 0.73 | 0.83 | 0.99 | 0.92 | 0.88 | 0.61 |
| DeepSeek-V4-Pro | 0.98 | 0.71 | 0.97 | 1.00 | 0.93 | 0.93 | 0.76 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Level | Label | Active fraction | Mean degree | Hand-offs / agent | Group size | Length |
|---|---|---|---|---|---|---|
| 1 | Light | 0.22–0.33 | 2.2 | 1.2 | 3–10 / 6.0 | 10–10 / 10.0 |
| 2 | Steady | 0.17–0.33 | 2.1 | 1.3 | 3–10 / 6.0 | 11–13 / 12.2 |
| 3 | Busy | 0.19–0.26 | 3.5 | 2.2 | 5–10 / 7.5 | 14–16 / 14.3 |
| 4 | Dense | 0.22–0.28 | 3.5 | 2.3 | 5–11 / 8.2 | 12–15 / 14.5 |
| 5 | Very dense | 0.16–0.28 | 3.6 | 2.2 | 6–11 / 8.5 | 13–16 / 14.7 |
| 6 | Peak | 0.36–0.40 | 4.1 | 4.6 | 7–12 / 9.5 | 20–20 / 20.0 |
| Domain | |||
|---|---|---|---|
| Healthcare-operations | 0.60 | 0.50 | 0.20 |
| Customer-support | 0.50 | 0.40 | 0.10 |
| Software-engineering | 0.50 | 0.30 | 0.10 |
| Home-living | 0.30 | 0.30 | 0.10 |
| Social-gatherings | 0.20 | 0.20 | 0.10 |
| Personal-productivity | 0.20 | 0.10 | 0.10 |
| Hop | Artifacts | Goal verbatim | Instruction verbatim | Drop rate |
|---|---|---|---|---|
| 1 | 50 | 80% | 72% | 32% |
| 2 | 14 | 64% | 79% | 65% |
| 3 | 5 | 60% | 80% | 64% |
| Target model | |||||||
|---|---|---|---|---|---|---|---|
| Main results: goals sampled across all categories (Table 1 ) | |||||||
| GPT-5.6 Luna | 0.38 | 0.32 | 0.33 | 0.80 | 0.57 | 0.40 | 0.24 |
| Kimi-K2.6 | 0.47 | 0.37 | 0.50 | 0.78 | 0.67 | 0.61 | 0.44 |
| GPT-OSS-120B | 0.85 | 0.73 | 0.83 | 0.99 | 0.92 | 0.88 | 0.61 |
| DeepSeek-V4-Pro | 0.98 | 0.71 | 0.97 | 1.00 | 0.93 | 0.93 | 0.76 |
| Security-critical goals only ( action category) | |||||||
| Target model | Jobs | Target cost (USD) | Outcome |
|---|---|---|---|
| GPT-OSS-120B | 109 | 1.71 | stopping rule met |
| Kimi-K2.6 | 372 | 33.18 | stopping rule met |
| DeepSeek-V4-Pro | 932 | 89.44 | stopping rule met |
| GPT-5.6 Luna | 4,002 | 198.58 | stopping rule met |
| Grok 4.6 | 571 | 58.97 | stopped for cost |
| Reproduction number | ||||
|---|---|---|---|---|
| Target model | Steps 0–15 ( ) | Steps 16–30 | Dispersion | Generation interval |
| GPT-5.6 Luna | 1.32 (19) | 0.94 | 9 | |
| Kimi-K2.6 | 2.12 (25) | 0.41 | 0.37 | 10 |
| GPT-OSS-120B | 1.57 (30) | 0.57 | 1.18 | 10 |
| DeepSeek-V4-Pro | 1.59 (29) | 0.68 | 0.62 | 10 |
| All models | 1.66 | — | 0.79 | 10 |
| Test universes | Large universes | |||
|---|---|---|---|---|
| Target model | Upload rate | Docs/universe | Upload rate | Docs/universe |
| GPT-5.6 Luna | 0.96 | 4.2 | 1.00 | 29.0 |
| Kimi-K2.6 | 0.97 | 7.0 | 0.96 | 42.7 |
| GPT-OSS-120B | 0.97 | 13.7 | 0.98 | 73.0 |
| DeepSeek-V4-Pro | 1.00 | 15.8 | 1.00 | 63.0 |
| Memory lists | Flagged | Score range | |
|---|---|---|---|
| Infected (should be flagged) | |||
| Main template, judge-verified | 104 | 104 | |
| Command and URL lines removed | 5 | 5 | 0.97–0.98 |
| Held-out template variant | 5 | 5 | |
| Instructions paraphrased by an LLM | 12 | 12 | 0.96–0.99 |
| Twelve distinct template styles | 12 | 11 | 0.31–0.94 |
| Study | State carrying the attack | Continuation path | Goal evaluation |
|---|---|---|---|
| Hidden in Memory ( Pulipaka et al., 2026 ) | Private agent memory | The same assistant uses the poisoned memory in later conversations; transfer to another assistant is not evaluated. | Reusable template across many goals. |
| MURMUR ( Patlan et al., 2025 ) | Persistent state shared by users of one collaborative agent | A later user interacts with that same agent and is affected by its shared state. | Attacker-specified actions. |
| Prompt Infection ( Lee & Tiwari, 2024 ; Yu et al., 2024 ) | Agent messages and active context | Infected agents pass instructions directly to other agents in a connected multi-agent system. | Attack-specific payloads. |
| AI Worm ( Cohen et al., 2025 ) | Email indexed by an assistant’s RAG system | An assistant reproduces the prompt in outgoing email, which can enter another assistant’s email and RAG pipeline. | Predefined malicious effects. |
| AgentWorm ( Zhang et al., 2026b ) | Agent configuration and other persistent files | An infected agent autonomously reaches peers through messaging and agent ecosystem channels. | Three payload types. |
| Autonomous LLM Agent Worms ( Zha & Wang, 2026 ) | File-backed workspace and memory state | Persistent content re-enters execution and drives autonomous cross-platform transmission. | Selected malicious actions. |