Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning
Authors: Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu
Organizations: Sydney AI Centre, The University of Sydney · Australian Institute for Machine Learning, Adelaide University · Australian Artificial Intelligence Institute, University of Technology Sydney · Wuhan University · School of Mathematics and Statistics, The University of Melbourne
Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counterfactual scenarios reveal how beliefs change as the latent concept of memory shifts under hypothetical interventions; thus, measuring the resulting concept drift provides a signal for distinguishing faithful versus false memory. Such concepts can be estimated from agent hidden states before answer generation, avoiding the need for reward design or answer sampling. Empirical experiments reveal that simply monitoring answers often fails to detect false memory, while FAME achieves AUROCs of 76.2% - 96.7% across false-memory settings, and outperforms the best baseline by 3.4% - 23.3% across realistic benchmarks, spanning math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard). We further release corresponding counterfactual templates and facilitate future research on false memory.
Figures & tables
Figure 1: Illustration of FAME: It evaluates false memory through counterfactual reasoning by varying memory factors to observe concept drift. The counterfactual query and memory jointly establish boundaries for identifying false memory.
Figure 2: Causal graph of agent belief update.
Figure 3: False-memory taxonomy in practice and examples of constructing counterfactual queries.
Table 4
Figure 6: Small drift yet crossing the boundary still signals false memory.
GSM-Symbolic
Gitchameleon
BigBench-Hard
Method
SK
SR
ES
KC
SK
SR
Input similarity
0.464
0.705
0.797
0.621
0.230
0.636
Surface
0.553
0.602
0.826
0.564
0.437
0.634
Logit confidence
0.530
0.372
0.539
0.452
0.669
0.762
Task vector
0.485
0.289
0.651
0.564
0.475
0.417
Function vector
0.595
0.660
0.781
0.665
0.576
0.837
Table 2: Effectiveness of false-memory evaluation in realistic benchmarks (AUROC ↑ ). Best results are in bold .
Figure 7: Two case studies of false memory in practice that FAME successfully evaluates.
Figure 8: Memory size vs. false-memory rate (spurious-keeping, GSM) and vs. retention rate (knowledge conflict, Git).
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Twin network for the belief update mechanism.
Figure 10: Layer sweep results of Llama-3.2-3B-Instruct.
Figure 11: Case studies of false memory caused by spurious correlation (GSM Benchmark).
Figure 12: Case study of concept changes with different memory sizes.
Figure 13: Case study of concept changes with different memory sizes.
Method
AUROC
Diff.
Spurious Removal (Robustness Expectation)
FAME (Ours)
0.756 ± 0.096
–
w/o robustness expectation
0.559 ± 0.155
↓ 0.197
cross-regime expectation swap
0.553 ± 0.138
↓ 0.203
random-direction gˉO
0.503 ± 0.192
↓ 0.254
Spurious Keeping (Adaptation Expectation)
Appendix
Table 3: Ablation study, AUROC (mean ± std over 10 random 50% subsamples of the top-70% template pool) on GSM-Symbolic with Llama-3.2-3B-Instruct . Diff. is relative to FAME’s own mean within each regime.
Method
AUROC
Diff.
FAME (Ours)
0.777 ± 0.040
–
reference cloud → average point
0.676 ± 0.081
− 0.101
non-filtering → filtering
0.756 ± 0.096
+ 0.002
Appendix
Table 4: Robustness of reference cloud and readout reference signal (on GSM, measured by AUROC).
Figure 14: Layer sweep for selecting the concept-extraction layer. Left: Llama-8B, Mid: Mistral-7B, Right: Qwen2.5-3B.
Method
SK
SR
ES
Input similarity
0.331 ± 0.118
0.652 ± 0.074
0.723 ± 0.026
Surface
0.412 ± 0.095
0.546 ± 0.127
0.819 ± 0.008
Logit confidence
0.717 ± 0.104
0.427 ± 0.145
0.534 ± 0.033
Task vector
0.521 ± 0.072
0.193 ± 0.056
0.622 ± 0.036
FAME (Ours)
0.786 ± 0.091
0.914 ± 0.047
0.848 ± 0.006
Appendix
Table 5: Effectiveness of FAME using Llama-3.1-8B-Instruct on GSM benchmark, measured by AUROC.
Method
SK
SR
ES
Input similarity
0.564 ± 0.048
0.670 ± 0.050
0.645 ± 0.017
Surface
0.563 ± 0.051
0.659 ± 0.041
0.810 ± 0.021
Logit confidence
0.478 ± 0.055
0.292 ± 0.060
0.658 ± 0.020
Task vector
0.505 ± 0.053
0.445 ± 0.076
0.761 ± 0.012
FAME (Ours)
0.623 ± 0.034
0.791 ± 0.032
0.843 ± 0.008
Appendix
Table 6: Effectiveness of FAME using Mistral-7B-Instruct on GSM benchmark, measured by AUROC.
Method
SK
SR
ES
Input similarity
0.394 ± 0.043
0.634 ± 0.032
0.663 ± 0.065
Surface
0.595 ± 0.094
0.470 ± 0.072
0.825 ± 0.027
Logit confidence
0.400 ± 0.043
0.488 ± 0.087
0.721 ± 0.058
Task vector
0.514 ± 0.076
0.641 ± 0.032
0.274 ± 0.060
FAME (Ours)
0.776 ± 0.046
0.775 ± 0.028
0.850 ± 0.022
Appendix
Table 7: Effectiveness of FAME using Qwen2.5-3B-Instruct on GSM benchmark, measured by AUROC.
Figure 15: Examples of counterfactual templates for spurious keeping, spurious removal, and environment shift in GSM-Symbolic.
Figure 16: Examples of counterfactual templates for knowledge conflict in GitChameleon.
Figure 17: Examples of counterfactual templates for spurious keeping and spurious removal in BigBench-Hard.
Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures. We show that this assumption can fail systematically: across ALFWorld and HumanEval, agents store confident but incorrect interpretations of the task and continue acting on them across trials, even though the environment resets to the correct task each time. We call this failure mode memory confabulation and introduce the Reflection Repetition Rate (RRR), a log-based metric that detects repeated reliance on incorrect reflective content. Using RRR, we identify 16 frozen environments in ALFWorld, where 0 of 121 reflections mention the correct target object, and 4 analogous cases in HumanEval. Our mitigation replaces open-ended self-diagnosis with programmatic extraction of trajectory-level failure signals, increasing correct object mention from 0% to 86%, reducing RRR from 0.64 to 0.10, and solving 3 of 16 frozen ALFWorld environments, suggesting that reflective memory can reinforce false beliefs rather than correct them.
Prakhar Dixit, Sadia Kamal, Tim Oates
Department of Computer Science, University of Maryland Baltimore County, Baltimore, MD, USA.
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlooking how retrieved memories influence downstream reasoning and decision-making. To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. MemSyco-Bench measures when memory should influence a decision and how valid memory should be used. Specifically, it covers five tasks that assess whether agents can reject memory as factual evidence, respect its applicable scope, resolve conflicts between memory and objective evidence, track memory updates, and use valid memory for personalization. All related resources are collected for the community at https://github.com/XMUDeepLIT/MemSyco-Bench.
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.
Mingfei Lu, Mengjia Wu, Runsong Jia +2
Australian Artificial Intelligence Institute (AAII) University of Technology Sydney