cs.DCSep 29, 2026

When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents

Authors: Yuqiao Meng, Luoxi Tang, Yingxue Zhang, Yuchen Yang, Zhaohan Xi

Organizations: Binghamton University, State University of New York, Binghamton, NY, USA · The Pennsylvania State University, University Park, PA, USA

Abstract

Persistent memory helps LLM agents carry information across long interactions, but correct memory can still be used incorrectly when queries change or memory states evolve. Existing work mainly studies memory content errors or evaluates fixed test cases, leaving memory-use failures hard to discover systematically. We formulate this issue as a fuzzing problem and categorize such failures into query-related and memory-state failures. We then develop U-Fuzz, which starts from memory checkpoints as test seeds, mutates queries or memory states under explicit mutation obligations, validates each mutant, and uses observed memory behavior to guide iterative testing while keeping failure labels outside the search. We evaluate U-Fuzz across several memory systems against diverse fuzzing baselines, and further test an output-only setting with API-based LLMs where memory retrieval is hidden. Across these settings, U-Fuzz consistently uncovers more confirmed memory-use failures, showing that its search remains effective across different memory architectures and even when only final responses are observable.

Figures & tables

Explore similar work

CardsList
  1. MemFail: Stress-Testing Failure Modes of LLM Memory Systems

    May 26, 2026Ishir Garg, Neel Kolhe, Dawn Song +1Large Language Model MemoryLarge Language Models Fail

  2. From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

    Sep 28, 2026Mingxi Zou, Langzhang Liang, Zhuo Wang +3Large Language Model AgentsHarms