Persistent memory can improve personalization in LLM agents but can also induce sycophancy and cross-domain leakage. We distinguish two governance decisions: admission, which determines what recalled information enters the working context, and presentation, which determines how admitted information is expressed. We implement two inference-time designs without retraining: factor-compiled admission (FC), which assesses whole memory entries, and permission-semantic admission (PS), which decomposes entries into typed units; both translate adjudicated attributes into eligibility decisions via deterministic policies. We evaluate on a four-backbone development suite and an external benchmark with four tasks of 300 samples each. Relative to verbatim injection, FC and PS reduce pooled judge-assessed failure rates on the external benchmark by 6.7 and 8.8 percentage points (p = 2.7e-7 and 4.1e-12), and development-set cross-domain leakage falls by up to 29.5 percentage points. A query-conditioned gating baseline shows no significant change in objective-fact failure or pooled failure. Under matched admission budgets, PS outperforms random and relevance-based selection on external objective-fact judgment after Holm correction. Holding presentation fixed, tightening admission cuts cross-domain failure by a further 17.5 percentage points (p = 1.6e-4); in contrast, no comparison between two renderings of identical adjudicated outputs survives multiple-comparison correction. Both designs increase personalization failures, and PS misses the preregistered improvement and personalization-preservation criteria. These results support evaluating admission and presentation separately: selection quality provides task-specific safety gains, while preserving beneficial memory use remains unresolved.
Figures & tables
Figure 1. Governance decomposed into two levers over one adjudication. Admission (upper half): a single batched adjudication assigns attributes to retrieved memories in situ; two admission designs (FC, PS) compile the adjudicated attributes to decide what enters the context, and excluded material is physically absent. Presentation (lower half): the same admitted verdicts are rendered under either sectioned form (PN vs. CH) —the verdict-shared comparison of Section 5.3 varies only this stage.
Backbone
Dim
Bare
PS+PN
p
PS+CH
p
MiniMax-M3
syc
31.1
27.0
.40
29.5
.83
MiniMax-M3
ben
16.5
18.7
.82
16.5
1.0
MiniMax-M3
xd
9.5
4.0
.052
4.0
.027
Qwen3.7-Max
syc
50.0
36.1
.0009
41.8
.053
Qwen3.7-Max
ben
0.0
0.0
1.0
0.0
1.0
Qwen3.7-Max
xd
44.0
14.5
<10−4
18.0
<10−4
Table 1. Development surface: failure rate (%) under verbatim injection (Bare) vs. PS+PN and PS+CH; paired exact McNemar, raw p . Dimensions: syc = sycophancy, ben = beneficial use, xd = cross-domain leakage. Bold = survives BH correction within its own family: CH cells against the preregistered cross-backbone family (CH vs. Bare on syc/xd, four backbones); PN cells against a separate exploratory family.
Task
Bare
Gate
p
FC
p
PS
p
ofj
45.2
45.5
1.0
23.7
1.4×10−13
20.4
7.6×10−19
csc
33.3
22.7
.0004
27.0
.042
25.3
.0049
mec
91.7
85.7
.0003
85.0
8.8×10−5
84.0
3.4×10−5
pm
14.7
22.7
.0012 †
22.3
.0022 †
20.0
.0365 †
pooled ( n=1,199 )
46.2
44.1
.08
39.5
2.7×10−7
37.4
4.1×10−12
Table 2. External benchmark: judged failure rate (%) by task, n=300 per cell (one ofj cell 299 after pairwise exclusion). Paired exact McNemar vs. Bare. Gate = our reimplementation of the query-conditioned gating policy of MemGate ( Zhang et al., 2026a ) on the same substrate (memories admitted by query-conditioned relevance, no adjudication). pm failure = appropriate stored preference not applied (lower is better). † marks a failure-rate increase relative to Bare (Section 5.5 ). Task abbreviations: ofj = objective fact judgment, csc = contextual scope control, mec = memory–evidence conflict, pm = personalized memory use.
Task
Metric
PN
CH
p
ofj
accuracy
.737
.777
.043
ofj
misconception
.220
.190
.12
csc
accuracy
.170
.160
.76
csc
scope violation
.253
.237
.60
mec
misled
.833
.823
.68
mec
accuracy
.003
.007
1.0
Table 3. Verdict-shared rendering experiment: identical admission decisions, two sectioned renderings (PN vs. CH). Paired exact McNemar; n=300 per row.
Task
NoMem
PS+CH
pm: accuracy / pref. used
1.0 / 2.0
60.3 / 80.7
csc: accuracy / violation
0.0 / 0.0
16.0 / 23.7
mec: misled
0.0
82.3
Table 4. No-memory ablation compared with the PS+CH arm. Values are percentages. Removing memory sharply reduces personalization accuracy and preference use, while yielding zero observed scope violations and misleading responses; this ablation removes both beneficial and conflicting content (Section 5.4 ).
Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions. However, when stored user information is reintroduced into the model context, it can also influence responses in inappropriate or unrelated settings. We study two such failure modes in memory-augmented LLMs: cross-domain leakage, where memories from one life domain affect responses in another, and memory-induced sycophancy, where stored user beliefs make models more likely to agree with the user rather than respond truthfully. We apply a simple inference-time modification to how memories are presented to the model, without changing the model or the memory contents. Across seven models on PersistBench, we compare the commonly used all-in context format, where memories are injected as an unstructured list, with structured formats that partition memories by domain. This simple modification consistently reduces cross-domain leakage while preserving utility, with our strongest method reducing leakage by 8.8% on average relative to the baseline.
Hakeem Hannoon, Andrew Zhao, Mihir Narayan +2
University of Saskatchewan · University of California–Irvine · University of Wisconsin–Madison +2
Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment wastes limited memory budget on transient interactions while failing to preserve critical context for long horizon tasks. To address this gap, we investigate an underexplored question: can LLM based memory systems learn personalized memory policies? We introduce PerMemBench, the first benchmark for evaluating personalized memory systems, featuring multi year, multi domain interaction histories across diverse user personas. We further present the first empirical study of memory personalization, proposing session level storage gating, a lightweight framework that selectively bypasses memory operations for transient sessions. Our study confirms that personalization yields substantial retention gains under perfect gating, yet reveals that accurate gating remains an open and critical challenge.
Long-term memory systems allow LLM agents to preserve information beyond a single context window, but most systems focus on storing and retrieving facts after extraction, leaving the write decision under-specified. What deserves memory can depend on the user's current task, topic, activity, or interaction partner, while uniform extraction applies one notion of importance across these different situations. We formulate this challenge as preference-conditioned write control and introduce AdaMem, which uses adaptive natural-language Memory Policies to personalize what an agent writes to memory. Each policy represents the user's memory preference for a particular interaction context, is updated from periodic feedback, and controls subsequent memory writing. We evaluate this loop in AdaMem-Bench, which assigns different memory preferences to six concurrent interaction personas across five ten-week stories. Across two extraction models and two feedback modes, AdaMem improves average QA accuracy over Mem0 from 80.0% to 84.35% while reducing persistent memory by 9.27%. Our analyses show that explicit feedback helps models learn better memory policies, but current models still struggle to translate those policies into reliably selective writing behavior. AdaMem thus demonstrates the promise of adaptive write control while exposing policy execution as a central limitation of current memory agents. Our code is publicly available: https://github.com/galaxyChen/AdaMem