Large language models (LLMs) have become the foundation of personalized assistants, but maintaining persistent user memory across long-term interactions remains challenging. Existing memory systems often focus on storage, retrieval, or consolidation, while memory writing remains less controlled: transient requests, duplicate statements, and outdated user states may enter memory and later be retrieved for personalization. In this paper, we present AMU: Admission and Memory Update for Personalized Conversations, an SLM-guided (Small language model guided) structured framework for writing-time memory control. AMU uses structured memory filtering to decide what should enter memory and SLM-guided storage management to determine whether an admitted record should be stored separately, discarded as a duplicate, or fused as an update. We evaluate AMU in a controlled memory writing and retrieval setting. Experimental results show that AMU maintains cleaner and more retrievable personalized memories.
Figures & tables
Figure 1: Motivation of writing-time memory control. Directly storing user utterances can preserve transient, redundant, or outdated memories, while AMU keeps stable and reusable memories for cleaner personalization.
Figure 2: Overview of AMU. Structured fields guide memory admission and retrieval, while the SLM decides whether an admitted record is duplicate, update, or separate before storage.
Metric
Validity (%)
Action label validity
89.3
Query-memory alignment
94.1
Table 1: Manual validation of silver annotations. Reported values are mean validity rates across independent annotations.
Method
P@3
R@3
F1@3
HIT@3
Red.@3 ↓
Full Comp.
35.42
55.50
41.30
57.00
12.08
Sliding Window
24.42
29.50
25.87
29.50
11.75
Mem0
34.08
57.50
40.88
56.00
3.63
A-MAC
30.36
58.76
37.96
54.50
3.18
A-MEM
35.13
56.90
40.36
56.50
2.88
AMU(Ours)
37.33
63.50
44.47
57.50
2.83
Table 2: Main retrieval results. Red.@3 denotes Redundancy@3, where lower values are better.
Variant
F1@3
Red.@3 ↓
Act. Acc.
w/o Struct. Filtering
34.30
2.25
65.40
w/o SLM Storage Mgmt.
25.62
0.25
34.50
AMU(Ours)
44.47
2.83
67.70
Table 3: Component verification of AMU. Red.@3 denotes Redundancy@3, where lower values are better.
Size
F1@3
Red.@3 ↓
Act. Acc.
0.8B
13.90
0.31
26.40
2B
44.47
2.83
67.70
4B
41.78
3.08
81.00
9B
40.00
3.33
84.20
Table 4: Effect of SLM controller size.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Overall Score
No Memory
2.92
Full Compressed Memory
3.21
AMU
4.37
Appendix
Table 5: Small-scale downstream response evaluation with an LLM judge.
τs
F1@3
Red.@3 ↓
Acc.
Time/Turn ↓
0.40
43.68
2.75
66.80
327.17
0.60
44.47
2.83
67.70
231.34
0.80
45.82
4.50
65.80
176.74
Appendix
Table 6: Sensitivity to the writing-time threshold τs with τr=0.55 . Red.@3 denotes Redundancy@3, where lower values are better.
τr
P@3
R@3
F1@3
Red.@3 ↓
HIT@3
0.40
36.00
64.25
43.75
2.83
58.00
0.55
37.33
63.50
44.47
2.83
57.50
0.70
45.33
51.25
46.33
2.00
44.50
Appendix
Table 7: Sensitivity to the retrieval-time threshold τr with τs=0.60 . Red.@3 denotes Redundancy@3, where lower values are better.
Embedding Model
P@3
R@3
F1@3
Red.@3 ↓
HIT@3
Action Acc.
Qwen3-Embedding-0.6B
37.33
63.50
44.47
2.83
57.50
67.70
BGE-M3
36.83
63.00
43.97
2.92
57.00
67.10
Multilingual-E5-Large
36.58
62.75
43.68
2.75
56.50
66.90
Appendix
Table 8: Additional analysis of embedding models, SLM backbones, and session mixing. Red.@3 and Red. denote Redundancy@3, where lower values are better.
Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment wastes limited memory budget on transient interactions while failing to preserve critical context for long horizon tasks. To address this gap, we investigate an underexplored question: can LLM based memory systems learn personalized memory policies? We introduce PerMemBench, the first benchmark for evaluating personalized memory systems, featuring multi year, multi domain interaction histories across diverse user personas. We further present the first empirical study of memory personalization, proposing session level storage gating, a lightweight framework that selectively bypasses memory operations for transient sessions. Our study confirms that personalization yields substantial retention gains under perfect gating, yet reveals that accurate gating remains an open and critical challenge.
Long-term memory systems allow LLM agents to preserve information beyond a single context window, but most systems focus on storing and retrieving facts after extraction, leaving the write decision under-specified. What deserves memory can depend on the user's current task, topic, activity, or interaction partner, while uniform extraction applies one notion of importance across these different situations. We formulate this challenge as preference-conditioned write control and introduce AdaMem, which uses adaptive natural-language Memory Policies to personalize what an agent writes to memory. Each policy represents the user's memory preference for a particular interaction context, is updated from periodic feedback, and controls subsequent memory writing. We evaluate this loop in AdaMem-Bench, which assigns different memory preferences to six concurrent interaction personas across five ten-week stories. Across two extraction models and two feedback modes, AdaMem improves average QA accuracy over Mem0 from 80.0% to 84.35% while reducing persistent memory by 9.27%. Our analyses show that explicit feedback helps models learn better memory policies, but current models still struggle to translate those policies into reliably selective writing behavior. AdaMem thus demonstrates the promise of adaptive write control while exposing policy execution as a central limitation of current memory agents. Our code is publicly available: https://github.com/galaxyChen/AdaMem
Personalizing large language models (LLMs) requires encoding long-term, user-specific behavioral patterns in a way that is computationally efficient, scalable, and compatible with a frozen base model. We present Latent Personal Memory (LPM), a scalable framework that represents user-specific history as a compact, persistent matrix of N latent slots, that are interpretable. A shared cross-attention projection network maps these slots into dynamic, input-conditioned soft prompts that are prepended to the input of a frozen LLM. We evaluate LPM on PersonaMem v1 and LoCOMO benchmarks across Qwen3-1.7B, 4B, and 8B backbones. Results demonstrate that LPM outperforms LoRA and Prompt Tuning by up to 8.8% and 54.4% in overall accuracy respectively on PersonaMem v1, while reducing KV-cache usage by over 64x. On LoCoMo, LPM matches LoRA accuracy with 120x fewer trainable parameters. We also show that the efficiency of LPM grows with context length and outperforms full-context at 128K context length.