Persistent memory helps LLM agents carry information across long interactions, but correct memory can still be used incorrectly when queries change or memory states evolve. Existing work mainly studies memory content errors or evaluates fixed test cases, leaving memory-use failures hard to discover systematically. We formulate this issue as a fuzzing problem and categorize such failures into query-related and memory-state failures. We then develop U-Fuzz, which starts from memory checkpoints as test seeds, mutates queries or memory states under explicit mutation obligations, validates each mutant, and uses observed memory behavior to guide iterative testing while keeping failure labels outside the search. We evaluate U-Fuzz across several memory systems against diverse fuzzing baselines, and further test an output-only setting with API-based LLMs where memory retrieval is hidden. Across these settings, U-Fuzz consistently uncovers more confirmed memory-use failures, showing that its search remains effective across different memory architectures and even when only final responses are observable.
Figures & tables
Figure 1: Illustrations of two memory-use failures: (1) paraphrasing a query changes the evidence retrieved from the same memory state; (2) after a memory update, stale evidence is ranked above the current value.
Figure 2: Overview of U-Fuzz . U-Fuzz builds seeds and references from memory checkpoints, generates validated query and memory-state mutations, and executes each parent-mutant pair to detect memory-use failures. Retrieval coverage and divergence guide later fuzzing rounds, while failure labels remain outside the fuzzing search.
Method
Mem0
A-Mem
Graphiti
MemOS
UF@ B
Cov@ B
UF@ B
Cov@ B
UF@ B
Cov@ B
UF@ B
Cov@ B
LoCoMo
Random Mutation
12.7 ±1.5
0.40 ±.02
21.7 ±5.1
0.24 ±.08
18.7 ±3.2
0.27 ±.12
15.0 ±4.6
0.34 ±.10
Unguided LLM
22.0 ±2.0
0.47 ±.03
29.7 ±3.8
0.39 ±.08
27.3 ±4.0
0.35 ±.10
26.0 ±3.0
0.45 ±.12
LLM-as-Judge
23.3 ±2.5
0.41 ±.01
35.7 ±2.3
0.38 ±.04
23.3 ±1.5
0.34 ±.02
25.3 ±0.6
0.45 ±.02
Coverage-Guided
27.7 ±4.0
0.59 ±.03
31.3 ±2.5
0.54 ±.05
30.3 ±3.5
0.47 ±.04
30.0 ±1.0
0.65 ±.02
U-Fuzz-Q
33.0 ±2.6
0.57 ±.07
57.0 ±1.0
0.51 ±.03
39.7 ±2.9
0.42 ±.05
41.3 ±2.5
0.57 ±.03
Table 1: Performance under a fixed execution budget B ( B=8000 for LoCoMo and B=4000 for LongMemEval-S). Results report mean and standard deviation over three runs. Appendix C also provides some case studies of found memory-use failures.
Figure 3: Fuzzing progress as the execution budget (x-axis) increases on LoCoMo. Each column corresponds to one memory system. Shaded regions indicate one standard deviation over three runs.
Figure 4: Component ablation on LoCoMo. Each bar reports the change in UF@ B or Cov@ B after removing one component from the full U-Fuzz . Mutation ablations remove one query- or memory-state operator, while Feedback ablations remove one search-feedback indicator. Error bars show one standard deviation over three runs.
Figure 5: Failure discovery under output-only observability on LoCoMo. Each bar reports UF@ BB=2000 , wherein we prompt LLMs to output reasoning and responses. Error bars indicate one standard deviation over three runs.
Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empirical work has been done to understand the specific failure modes and design choices that these systems present. Existing benchmarks report aggregate question-answering accuracy and treat memory systems as black boxes, making it impossible to attribute an incorrect answer to a particular failure mode of the system. We introduce MemFail, a diagnostic benchmark that isolates the failure modes of modern LLM memory systems. We begin by formalizing memory systems as the composition of three canonical operations -- summarization, storage, and retrieval -- and identify the potential failure modes induced by each. Based on these hypothesized failure modes, we construct five datasets spanning four tasks, each adversarially designed to test a specific operation of a memory system. Using these datasets, we evaluate four state-of-the-art memory systems on MemFail and demonstrate how MemFail can be used to empirically understand the tradeoffs induced by differences in memory system architectures.
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.
Mingxi Zou, Langzhang Liang, Zhuo Wang +3
Shanghai Academy of AI for Science (SAIS) · Fudan University · Monash University
Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, an external observation may be rewritten as apparent user history or workflow support, preserving an action trigger while erasing the low-trust source that should limit its authority. Existing prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy memory consolidation. We formalize this boundary and instantiate it as Provenance-Preserving Memory Fire wall (PPMF), a lightweight memory middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. In our schema-grounded evaluation with fixed risk policies, vulnerable consolidated memories reach up to 1.000 attack success rate(ASR); with intact platform-maintained provenance, confirmation, and risk labels, no evaluated unauthorized high-risk action passes the PPMF gate while confirmed benign actions and targeted low-risk memory use remain executable.