Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning
Authors: Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu
Organizations: Sydney AI Centre, The University of Sydney · Australian Institute for Machine Learning, Adelaide University · Australian Artificial Intelligence Institute, University of Technology Sydney · Wuhan University · School of Mathematics and Statistics, The University of Melbourne
Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counterfactual scenarios reveal how beliefs change as the latent concept of memory shifts under hypothetical interventions; thus, measuring the resulting concept drift provides a signal for distinguishing faithful versus false memory. Such concepts can be estimated from agent hidden states before answer generation, avoiding the need for reward design or answer sampling. Empirical experiments reveal that simply monitoring answers often fails to detect false memory, while FAME achieves AUROCs of 76.2% - 96.7% across false-memory settings, and outperforms the best baseline by 3.4% - 23.3% across realistic benchmarks, spanning math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard). We further release corresponding counterfactual templates and facilitate future research on false memory.
Figures & tables
Figure 1: Illustration of FAME: It evaluates false memory through counterfactual reasoning by varying memory factors to observe concept drift. The counterfactual query and memory jointly establish boundaries for identifying false memory.
Figure 2: Causal graph of agent belief update.
Figure 3: False-memory taxonomy in practice and examples of constructing counterfactual queries.
Table 4
Figure 6: Small drift yet crossing the boundary still signals false memory.
GSM-Symbolic
Gitchameleon
BigBench-Hard
Method
SK
SR
ES
KC
SK
SR
Input similarity
0.464
0.705
0.797
0.621
0.230
0.636
Surface
0.553
0.602
0.826
0.564
0.437
0.634
Logit confidence
0.530
0.372
0.539
0.452
0.669
0.762
Task vector
0.485
0.289
0.651
0.564
0.475
0.417
Function vector
0.595
0.660
0.781
0.665
0.576
0.837
Table 2: Effectiveness of false-memory evaluation in realistic benchmarks (AUROC ↑ ). Best results are in bold .
Figure 7: Two case studies of false memory in practice that FAME successfully evaluates.
Figure 8: Memory size vs. false-memory rate (spurious-keeping, GSM) and vs. retention rate (knowledge conflict, Git).
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Twin network for the belief update mechanism.
Figure 10: Layer sweep results of Llama-3.2-3B-Instruct.
Figure 11: Case studies of false memory caused by spurious correlation (GSM Benchmark).
Figure 12: Case study of concept changes with different memory sizes.
Figure 13: Case study of concept changes with different memory sizes.
Method
AUROC
Diff.
Spurious Removal (Robustness Expectation)
FAME (Ours)
0.756 ± 0.096
–
w/o robustness expectation
0.559 ± 0.155
↓ 0.197
cross-regime expectation swap
0.553 ± 0.138
↓ 0.203
random-direction gˉO
0.503 ± 0.192
↓ 0.254
Spurious Keeping (Adaptation Expectation)
Appendix
Table 3: Ablation study, AUROC (mean ± std over 10 random 50% subsamples of the top-70% template pool) on GSM-Symbolic with Llama-3.2-3B-Instruct . Diff. is relative to FAME’s own mean within each regime.
Method
AUROC
Diff.
FAME (Ours)
0.777 ± 0.040
–
reference cloud → average point
0.676 ± 0.081
− 0.101
non-filtering → filtering
0.756 ± 0.096
+ 0.002
Appendix
Table 4: Robustness of reference cloud and readout reference signal (on GSM, measured by AUROC).
Figure 14: Layer sweep for selecting the concept-extraction layer. Left: Llama-8B, Mid: Mistral-7B, Right: Qwen2.5-3B.
Method
SK
SR
ES
Input similarity
0.331 ± 0.118
0.652 ± 0.074
0.723 ± 0.026
Surface
0.412 ± 0.095
0.546 ± 0.127
0.819 ± 0.008
Logit confidence
0.717 ± 0.104
0.427 ± 0.145
0.534 ± 0.033
Task vector
0.521 ± 0.072
0.193 ± 0.056
0.622 ± 0.036
FAME (Ours)
0.786 ± 0.091
0.914 ± 0.047
0.848 ± 0.006
Appendix
Table 5: Effectiveness of FAME using Llama-3.1-8B-Instruct on GSM benchmark, measured by AUROC.
Method
SK
SR
ES
Input similarity
0.564 ± 0.048
0.670 ± 0.050
0.645 ± 0.017
Surface
0.563 ± 0.051
0.659 ± 0.041
0.810 ± 0.021
Logit confidence
0.478 ± 0.055
0.292 ± 0.060
0.658 ± 0.020
Task vector
0.505 ± 0.053
0.445 ± 0.076
0.761 ± 0.012
FAME (Ours)
0.623 ± 0.034
0.791 ± 0.032
0.843 ± 0.008
Appendix
Table 6: Effectiveness of FAME using Mistral-7B-Instruct on GSM benchmark, measured by AUROC.
Method
SK
SR
ES
Input similarity
0.394 ± 0.043
0.634 ± 0.032
0.663 ± 0.065
Surface
0.595 ± 0.094
0.470 ± 0.072
0.825 ± 0.027
Logit confidence
0.400 ± 0.043
0.488 ± 0.087
0.721 ± 0.058
Task vector
0.514 ± 0.076
0.641 ± 0.032
0.274 ± 0.060
FAME (Ours)
0.776 ± 0.046
0.775 ± 0.028
0.850 ± 0.022
Appendix
Table 7: Effectiveness of FAME using Qwen2.5-3B-Instruct on GSM benchmark, measured by AUROC.
Figure 15: Examples of counterfactual templates for spurious keeping, spurious removal, and environment shift in GSM-Symbolic.
Figure 16: Examples of counterfactual templates for knowledge conflict in GitChameleon.
Figure 17: Examples of counterfactual templates for spurious keeping and spurious removal in BigBench-Hard.