cs.AISep 30, 2026

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning

Authors: Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu

Organizations: Sydney AI Centre, The University of Sydney · Australian Institute for Machine Learning, Adelaide University · Australian Artificial Intelligence Institute, University of Technology Sydney · Wuhan University · School of Mathematics and Statistics, The University of Melbourne

Abstract

Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counterfactual scenarios reveal how beliefs change as the latent concept of memory shifts under hypothetical interventions; thus, measuring the resulting concept drift provides a signal for distinguishing faithful versus false memory. Such concepts can be estimated from agent hidden states before answer generation, avoiding the need for reward design or answer sampling. Empirical experiments reveal that simply monitoring answers often fails to detect false memory, while FAME achieves AUROCs of 76.2% - 96.7% across false-memory settings, and outperforms the best baseline by 3.4% - 23.3% across realistic benchmarks, spanning math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard). We further release corresponding counterfactual templates and facilitate future research on false memory.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Honest Lying: Understanding Memory Confabulation in Reflexive Agents

    May 28, 2026Prakhar Dixit, Sadia Kamal, Tim OatesReflectionsConfabulation

  2. MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

    Jul 1, 2026Zhishang Xiang, Zerui Chen, Yunbo Tang +5SycophancyArtificial Intelligence Agents

  3. Remember Before You're Asked: MemDream for Self-Probing Memory Evolution

    Sep 28, 2026Mingfei Lu, Mengjia Wu, Runsong Jia +2Large Language Model AgentsDream