From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents
Organizations: Shanghai Academy of AI for Science (SAIS) · Fudan University · Monash University
Abstract
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.
Figures & tables
| Backbone | Attack | AgentDojo | -bench | ||||||
|---|---|---|---|---|---|---|---|---|---|
| CMR | ASR (%) | BUD (pp) | (%) | CMR | ASR (%) | BUD (pp) | (%) | ||
| GPT-4o | General attack baselines | ||||||||
| AgentPoison | |||||||||
| MINJA | |||||||||
| InjecMEM | |||||||||
| MemPoison | |||||||||
| GPT-4o | Llama-3.1-70B | |||
| Attack / comparison | AgentDojo | -bench | AgentDojo | -bench |
| (a) General attacks: common-support CMR identical six-contract support | ||||
| AgentPoison | ||||
| MINJA | ||||
| InjecMEM | ||||
| MemPoison | ||||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Attack variable | Reported insertion / retention path | Role of retrieval in the reported mechanism | Reported detectability / plausibility framing | Evaluation target |
|---|---|---|---|---|---|
| PoisonedRAG | Malicious passages | Retrieval-corpus or index content; agent-memory state is not the primary target | Central to the attack | Perplexity-based and duplicate-filter defenses evaluated | Attacker-chosen answer for an attacker-chosen target question |
| AgentPoison | Trigger-conditioned poisoned demonstrations | Poisoned long-term-memory or knowledge-base entries | Central in the reported retrieval-equipped configurations | Trigger coherence; benign-task behavior also measured | Triggered target action and end-to-end adversarial impact |
| MINJA | Query-induced malicious records with bridging steps | Records induced through query-only interaction and retained by the agent | Central to later activation; query-only describes the insertion route | Query-only interaction constraint | Targeted malicious reasoning on later victim queries |
| InjecMEM | Topic anchor paired with an instruction payload | A record induced in one interaction and retained through the memory path | Retriever-agnostic anchor designed for topic-conditioned later retrieval | Topic-conditioned retrieval | Pre-specified output for related queries |
| MemoryGraft | Poisoned successful experiences or procedure templates | Ingestion artifacts that induce persistent poisoned experience records | Experience retrieval is central | Successful-looking traces | Persistent behavioral drift on semantically similar tasks |
| MemPoison | Dialogue-delivered trigger–payload relation | A triggerable backdoor retained through selective memory processing | Depends on the memory extraction/retrieval pipeline | Selective-memory survival | Trigger-conditioned later responses |
| Memory | Write unit | Read interface | Factor representation | Defense-relevant property |
|---|---|---|---|---|
| Vector episodic | Raw episode or document chunk | Similarity retrieval | Implicit span-level fact | Preserves source spans, but over-retrieves plausible edits |
| Rolling summary | Session or task summary | Injected summary context | Compressed natural-language rule | Loses fine-grained provenance during compression |
| User profile | Structured or semi-structured profile field | Profile lookup / injection | Explicit preference, threshold, or binding | Rejects unsupported fields under strict schemas |
| Reflection store | Generated lesson or rule | Retrieved reflection | Reusable behavioral prior | Converts local facts into broad rules |
| Graph memory | Entity–relation edge | Graph neighborhood lookup | Explicit typed edge | Benefits from schema validation and signed edges |
| Hybrid memory | Retrieval plus summary/profile state | Mixed retrieval and injection | Implicit and explicit | Inherits provenance loss and structured validation hooks |
| Selection policy | CMR | ASR (%) |
|---|---|---|
| Random | 0.6830 | 57.55 |
| ASR-guided | 0.6958 | 75.30 |
| CMR-guided | 0.7830 | 71.28 |