Long-term memory enables LLM-based agents to retain and reuse information across tasks and sessions, supporting personalization and long-horizon interactions. However, persistent memories can also induce sycophancy, causing agents to over-align with users' historical beliefs even when they are inaccurate, outdated, or inconsistent with objective evidence. Existing mitigation methods assume that memory-induced sycophancy originates from biased or incorrect memories and attempt to reduce this risk by filtering such memories at different stages of the memory pipeline. However, in the real world, objective and correct memories can still induce sycophancy, and the same memory can warrant different influence across different contexts. To this end, we propose MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning. Specifically, MemAdapter consists of three components: (i) Counterfactual Induction, which leverages counterfactual reasoning to uncover the potential risk of retrieved memories; (ii) Context-Aware Reflection, which calibrates the inferential influence of each retrieved memory in light of the current task via self-reflection; and (iii) Evidence-Based Reasoning, which grounds the final response in appropriate evidence while preserving the legitimate influence of memory. Extensive experiments on three benchmarks demonstrate that MemAdapter consistently improves memory reliability across diverse scenarios. Our code is available at https://github.com/DEEP-JLU/MemAdapter.
Figures & tables
Figure 1: Two paradigms of memory-agent.
Figure 2: Preliminary study on memory-induced sycophancy. (a) Accuracy with and without retrieved memories, using only factually correct and task-relevant memories. (b) A fixed-memory case showing that the same memory may warrant different influence across contexts: it can support response personalization in one task but should not override task-specific evidence in another.
Figure 3: The overall pipeline of MemAdapter . It consists of three components: (i) Counterfactual Induction, which leverages counterfactual reasoning to uncover the potential risk of retrieved memories; (ii) Context-Aware Reflection, which calibrates the inferential influence of each retrieved memory in light of the current task via self-reflection; and (iii) Evidence-Based Reasoning, which grounds the final response in appropriate evidence while preserving the correct influence of memory.
Memory System
Method
MemSyco-Bench
PersistBench
MemTrapBench
When to Use Memory Acc. ↑
How to Use Memory Acc. ↑
Cross-domain FR@3 ↓
Sycophancy FR@3 ↓
Beneficial FR ↓
Reasoning Fixation Score ↑
Belief Distortion Score ↑
NaiveRAG
Direct Gen.
74.20
64.77
34.50
80.00
1.00
3.32
4.79
Anti-Syc.
89.78 (+15.58)
83.39 (+18.62)
37.00 (+2.50)
69.50 (-10.50)
4.00 (+3.00)
3.37 (+0.05)
4.84 (+0.05)
Self-ReC.
76.89 (+2.69)
86.31 (+21.54)
24.50 (-10.00)
75.00 (-5.00)
5.00 (+4.00)
3.96 (+0.64)
4.80 (+0.01)
Dyn. Part.
88.89 (+14.69)
84.77 (+20.00)
37.00 (+2.50)
76.00 (-4.00)
3.00 (+2.00)
3.43 (+0.11)
4.79 (+0.00)
MemGate
72.89 (-1.31)
77.38 (+12.61)
36.50 (+2.00)
82.00 (+2.00)
1.00 (+0.00)
3.36 (+0.04)
4.71 (-0.08)
Table 1: Main results across MemSyco-Bench, PersistBench, and MemTrapBench. The direct-generation row in each memory-system block denotes the corresponding memory system without an additional memory-use intervention. Parentheses report absolute differences relative to the direct-generation result. Blue–gray denotes improvement and red denotes degradation; darker blue–gray shades indicate larger improvements within the same memory-system block and metric. Underlined values denote the best result within each memory-system block for the corresponding metric.
Memory System
Method
Objective
Contextual
Memory–Evidence
Personalized
Valid Memory
Avg.
Fact Judgment
Scope Control
Conflict
Memory Use
Selection
Acc. ↑
Acc. ↑
Acc. ↑
Acc. ↑
Acc. ↑
Acc. ↑
GPT
NaiveRAG
Direct Gen.
74.00
92.31
99.67
75.67
85.71
85.48
MemAdapter
80.33 (+6.33)
97.00 (+4.69)
99.67 (+0.00)
70.67 (-5.00)
94.29 (+8.58)
88.58 (+3.10)
A-MEM
Direct Gen.
76.00
92.00
100.00
79.67
82.00
85.81
Table 2: Cross-backbone performance comparison on MemSyco-Bench with GPT-5.6-sol and Qwen3-8B. Direct Gen. denotes directly generating a response with the retrieved memories, without applying an additional memory-use intervention. Parentheses denote percentage-point changes relative to the corresponding direct-generation result.
Memory System
Ablation Variant
Objective Fact Judgment Acc. ↑
Contextual Scope Control Acc. ↑
Memory–Evidence Conflict Acc. ↑
Personalized Memory Use Acc. ↑
Valid Memory Selection Acc. ↑
Avg. Acc. ↑
NaiveRAG
Direct Gen.
59.33
79.00
84.28
49.00
78.29
70.25
B
69.33 (+10.00)
93.00 (+14.00)
98.00 (+13.72)
81.00 (+32.00)
80.57 (+2.28)
84.26 (+14.01)
B+R
75.67 (+16.34)
89.00 (+10.00)
99.33 (+15.05)
79.00 (+30.00)
83.14 (+4.85)
85.16 (+14.91)
MemAdapter
79.00 (+19.67)
94.00 (+15.00)
99.33 (+15.05)
81.00 (+32.00)
95.43 (+17.14)
89.94 (+19.69)
A-MEM
Direct Gen.
61.05
83.00
82.55
58.34
73.35
71.71
B
72.67 (+11.62)
93.33 (+10.33)
98.00 (+15.45)
78.67 (+20.33)
79.14 (+5.79)
84.19 (+12.48)
Table 3: Ablation results of MemAdapter on MemSyco-Bench. The five task-specific accuracy metrics and the average accuracy are reported. Parentheses denote changes relative to the corresponding direct-generation result. Within each memory system and each metric, darker blue backgrounds indicate larger improvements, whereas red backgrounds indicate degradations. Bold underlined values denote the best result within each memory-system block for the corresponding metric.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Aspect
Setting
Data composition
75 instances in total, with 25 instances derived from each of MemSyco-Bench, PersistBench, and MemTrapBench.
Sampling and construction
Candidate instances are first randomly sampled from each benchmark and manually reviewed against the study criteria. Each instance is retained when it already satisfies the controlled setting, or reconstructed when necessary to isolate the effect of memory use while preserving the underlying task scenario.
Instance
Each instance is represented as Ii=(qi,ci,mi,yi∗) , where qi is the current query, ci denotes task-specific evidence or constraints when applicable, mi is the supplied memory, and yi∗ is the gold answer.
Current task
The current query and available evidence are specified such that the gold answer can be determined from the current task itself without relying on the supplied memory.
Memory
The supplied memory is manually verified to be accurate with respect to the source scenario, benign, and semantically relevant to the current task. It may describe a genuine user preference, prior experience, strategy, or decision factor, but does not provide sufficient task-specific support to determine or override the gold answer.
Memory context
Each instance contains one to three mutually consistent memory units, avoiding noisy or contradictory retrieval as a confounding factor.
Appendix
Table 4: Controlled setup of the preliminary study.
Memory System
Method
Objective Fact Judgment Acc. ↑
Contextual Scope Control Acc. ↑
Memory–Evidence Conflict Acc. ↑
Personalized Memory Use Acc. ↑
Valid Memory Selection Acc. ↑
Avg. Acc. ↑
NaiveRAG
Direct Gen.
59.33
79.00
84.28
49.00
78.29
70.25
Anti-Syc.
76.67 (+17.34)
93.00 (+14.00)
99.67 (+15.39)
82.33 (+33.33)
84.29 (+6.00)
87.10 (+16.85)
Self-ReC.
73.00 (+13.67)
57.67 (-21.33)
100.00 (+15.72)
84.33 (+35.33)
88.00 (+9.71)
80.84 (+10.59)
Dyn. Part.
73.00 (+13.67)
93.67 (+14.67)
100.00 (+15.72)
81.67 (+32.67)
87.43 (+9.14)
87.16 (+16.91)
MemGate
75.67 (+16.34)
51.67 (-27.33)
91.33 (+7.05)
73.00 (+24.00)
81.14 (+2.85)
74.77 (+4.53)
MemAdapter
79.00 (+19.67)
94.00 (+15.00)
99.33 (+15.05)
81.00 (+32.00)
95.43 (+17.14)
89.94 (+19.69)
Appendix
Table 5: Detailed task-wise results on MemSyco-Bench. Direct Gen. denotes directly generating a response with the retrieved memories, without applying an additional memory-use intervention. Parentheses denote changes relative to the corresponding direct-generation result. Blue backgrounds indicate improvements, with darker shades representing larger improvements within the same memory system. Red backgrounds indicate degradations. Bold underlined values denote the best result within each memory-system block for the corresponding metric.
Memory System
Method
Task Boundary ↑
Cognitive Bias ↑
Trauma ↑
Safety ↑
Average ↑
NaiveRAG
Direct Gen.
3.40
3.24
4.85
4.74
4.06
Anti-Syc.
3.54 (+0.14)
3.21 (-0.03)
4.83 (-0.02)
4.84 (+0.10)
4.11 (+0.05)
Self-ReC.
4.45 (+1.05)
3.48 (+0.24)
4.92 (+0.07)
4.68 (-0.06)
4.38 (+0.32)
Dyn. Part.
3.50 (+0.10)
3.34 (+0.10)
4.86 (+0.01)
4.78 (+0.04)
4.12 (+0.06)
MemGate
3.45 (+0.05)
3.22 (-0.02)
4.86 (+0.01)
4.68 (-0.06)
4.05 (-0.01)
MemAdapter
4.85 (+1.45)
3.23 (-0.01)
4.83 (-0.02)
4.90 (+0.16)
4.45 (+0.40)
Appendix
Table 6: MemTrapBench results under the official scenario-level taxonomy. Direct Gen. generates responses directly from retrieved memories, without an additional memory-use intervention. Parentheses show changes from Direct Gen.; all scores are higher-is-better. Within each memory-system block and metric, darker blue indicates larger improvements and red indicates degradations. Bold underlined values mark the best results.
Memory System
Method
Cross-domain FR@1 ↓
Cross-domain FR@2 ↓
Cross-domain FR@3 ↓
Sycophancy FR@1 ↓
Sycophancy FR@2 ↓
Sycophancy FR@3 ↓
Beneficial FR ↓
NaiveRAG
Direct Gen.
23.00
30.00
34.50
61.50
71.50
80.00
1.00
Anti-Syc.
17.50 (-5.50)
30.50 (+0.50)
37.00 (+2.50)
45.00 (-16.50)
64.00 (-7.50)
69.50 (-10.50)
4.00 (+3.00)
Self-ReC.
15.00 (-8.00)
19.50 (-10.50)
24.50 (-10.00)
54.50 (-7.00)
69.00 (-2.50)
75.00 (-5.00)
5.00 (+4.00)
Dyn. Part.
16.00 (-7.00)
26.50 (-3.50)
37.00 (+2.50)
56.00 (-5.50)
66.50 (-5.00)
76.00 (-4.00)
3.00 (+2.00)
MemGate
25.00 (+2.00)
28.50 (-1.50)
36.50 (+2.00)
60.00 (-1.50)
76.00 (+4.50)
82.00 (+2.00)
1.00
MemAdapter
17.00 (-6.00)
23.50 (-6.50)
29.00 (-5.50)
35.00 (-26.50)
45.50 (-26.00)
53.50 (-26.50)
2.00 (+1.00)
Appendix
Table 7: Results on PersistBench. Direct Gen. generates responses from retrieved memories, without an additional memory-use intervention. Parentheses show changes from Direct Gen. All metrics are failure rates, so lower is better. Within each memory-system block, darker blue indicates larger improvements and red indicates degradations. Bold underlined values mark the best results.
Memory System
Ablation Variant
Objective Fact Judgment Acc. ↑
Contextual Scope Control Acc. ↑
Memory–Evidence Conflict Acc. ↑
Personalized Memory Use Acc. ↑
Valid Memory Selection Acc. ↑
Avg. Acc. ↑
NaiveRAG
Direct Gen.
59.33
79.00
84.28
49.00
78.29
70.25
B
69.33 (+10.00)
93.00 (+14.00)
98.00 (+13.72)
81.00 (+32.00)
80.57 (+2.28)
84.26 (+14.01)
B+R
75.67 (+16.34)
89.00 (+10.00)
99.33 (+15.05)
79.00 (+30.00)
83.14 (+4.85)
85.16 (+14.91)
MemAdapter
79.00 (+19.67)
94.00 (+15.00)
99.33 (+15.05)
81.00 (+32.00)
95.43 (+17.14)
89.94 (+19.69)
A-MEM
Direct Gen.
61.05
83.00
82.55
58.34
73.35
71.71
B
72.67 (+11.62)
93.33 (+10.33)
98.00 (+15.45)
78.67 (+20.33)
79.14 (+5.79)
84.19 (+12.48)
Appendix
Table 8: Ablation results of MemAdapter on MemSyco-Bench. The five task-specific accuracy metrics and the average accuracy are reported. Parentheses denote changes relative to the corresponding direct-generation result. Within each memory system and each metric, darker blue backgrounds indicate larger improvements, whereas red backgrounds indicate degradations. Bold underlined values denote the best result within each memory-system block for the corresponding metric.
Memory System
Method
Mean Time (s)
P50 Time (s)
P95 Time (s)
P99 Time (s)
A-MEM
Anti-Sycophancy
6.10
5.80
9.73
13.15
Dynamic Partition
6.67
6.38
10.31
14.08
MemGate
6.07
5.80
9.90
13.74
Self-ReCheck
6.77
6.47
10.91
13.99
MemAdapter
6.95
5.47
10.79
13.94
Mem0
Anti-Sycophancy
6.04
5.78
10.10
12.85
Appendix
Table 9: End-to-end runtime comparison of post-retrieval interventions on PersistBench. Mean, P50, P95, and P99 runtime are measured in seconds per sample under the corresponding execution settings. Bold underlined values denote the shortest runtime within each memory-system block for the corresponding metric.
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlooking how retrieved memories influence downstream reasoning and decision-making. To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. MemSyco-Bench measures when memory should influence a decision and how valid memory should be used. Specifically, it covers five tasks that assess whether agents can reject memory as factual evidence, respect its applicable scope, resolve conflicts between memory and objective evidence, track memory updates, and use valid memory for personalization. All related resources are collected for the community at https://github.com/XMUDeepLIT/MemSyco-Bench.
Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by systematically amplifying sycophancy, wherein models prioritize agreement with users over accuracy. We conduct the first systematic evaluation of this effect, introducing MIST: a benchmark of synthetically generated multi-turn conversations where users express plausible misconceptions in scientific, medical, and moral reasoning domains. Testing across three state-of-the-art memory systems and five model families reveals that memory amplifies sycophantic behavior across all conditions, with up to 25x higher sycophancy rates than in-context baselines. Error analyses suggest memory extraction as the primary culprit: lossy compression into discrete snippets encodes user misconceptions while discarding corrective context. Based on these results, we propose two lightweight mitigations that substantially reduce sycophancy while matching or exceeding memory systems at factual recall.
Shelly Bensal, Axel Magnuson, Aparna Balagopalan +1
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.
Mingfei Lu, Mengjia Wu, Runsong Jia +2
Australian Artificial Intelligence Institute (AAII) University of Technology Sydney