Long-term personalization requires language models to use interaction history to track users' preferences across sessions. Parametric memory encodes this interaction history into model parameters or adapters, reducing the need to include it in the inference context. However, independent context compilation leaves cross-session integration unspecified, while recurrent updates can attenuate earlier evidence. To address these challenges, we propose Dual-Path Parametric Memory (DPPM). Its Evidence path directly pools representations of the interaction history to preserve earlier evidence, while its Delta path sequentially updates an associative state to capture changes. Fusing both outputs produces history-conditioned LoRA adapters that combine evidence accumulation with ordered revision. Across multiple backbones, DPPM outperforms the evaluated baselines, achieving 54.22% on PersonaMem-v2 and 86.79% on PrefEval. These results suggest that DPPM provides a simple and effective design choice for cross-session personalized parametric memory.
Figures & tables
Figure 1: Illustrative comparison of cross-session memory. (a) Context compilation produces separate LoRA adapters. (b) Recurrent updates may dilute earlier evidence. (c) DPPM combines evidence accumulation and sequential revision to preserve dietary constraints while tracking preference changes.
Figure 2: Overall architecture of DPPM. A frozen compiler encodes the interaction history segment by segment, and DPPM consolidates the resulting representations through Evidence pooling and sequential Delta updates. Their calibrated outputs are fused and decoded into history-conditioned LoRA adapters.
Type
Method
PersonaMem-v2
PrefEval
Avg.
Overall
Self
Current
Overall
10
70
300
Reference
No Context
28.54
29.83
30.38
37.72
37.04
39.26
36.85
33.13
Gold State
61.08
67.44
67.44
93.52
92.78
94.07
93.70
77.30
Textual
Rolling Summary
29.16
30.93
31.12
40.49
38.70
40.00
42.78
34.83
LightMem
30.66
31.76
32.46
71.60
80.37
69.63
64.81
51.13
Mem0
30.46
31.93
33.06
75.49
77.96
76.11
72.41
52.98
Table 1: Main results on both benchmarks (accuracy, %). Self and Current select who=self and updated=False . PrefEval interval scores average three preference forms. Avg. averages the two Overall scores. Bold and underlining indicate the best and second-best scores, excluding the annotated Gold State reference. Metis uses its native Qwen3.5 backbone at the indicated size; the other methods use Qwen3-4B.
Backbone
Method
PersonaMem-v2
PrefEval
Overall
Self
Current
Overall
10
70
300
Gemma-2-2B IT
RPMem
29.17 ±0.34
30.28 ±0.40
30.88 ±0.44
75.01 ±0.92
73.07 ±1.70
75.70 ±1.03
76.26 ±1.30
Evidence only
38.20 ±0.23
41.34 ±0.24
44.50 ±0.24
79.81 ±0.58
77.52 ±0.71
81.33 ±1.19
80.59 ±1.25
Delta only
37.68 ±0.69
40.96 ±0.83
43.13 ±0.95
79.38 ±1.24
77.74 ±1.39
80.15 ±1.09
80.26 ±1.69
DPPM
40.31 ±0.54
43.71 ±0.63
46.77 ±0.79
79.20 ±0.76
78.56 ±0.44
79.56 ±0.59
79.48 ±2.60
Qwen3-4B Instruct-2507
RPMem
45.73 ±0.24
49.46 ±0.35
53.38 ±0.35
82.47 ±0.58
79.56 ±1.10
83.81 ±1.19
84.04 ±1.27
Table 2: Memory alternatives across three backbones (accuracy, %). All entries are five-run means, with sample standard deviations in subscripts. Best and second-best means are bold and underlined within each backbone.
Memory
Params
Persona
PrefEval
RPMem
524800
45.73 ±0.24
82.47 ±0.58
RPMem-matched
527364
45.49 ±0.42
82.56 ±0.90
DPPM
527364
54.22 ±0.48
86.79 ±0.68
Table 3: Scores are five-run Overall accuracy means (%), with sample standard deviations in subscripts. Persona denotes PersonaMem-v2.
Figure 3: Category-level comparison of RPMem and DPPM. Circles and diamonds show five-run mean accuracies, and the right-hand annotations give DPPM’s gain in percentage points. The panels cover all seven PersonaMem-v2 information types and all four held-out PrefEval topics. Axis ranges are shown separately for each benchmark.
Figure 4: Answer-stage efficiency on Qwen3-4B with one A800 (batch size 1). (a) Full-test accuracy versus weighted latency. (b) Peak GPU memory for paired histories of one length-selected question. Results use three-forward medians after one warmup, excluding compilation and retrieval; KV cache is disabled in (b).
Figure 5: Latent retention versus the number of subsequent distractor segments. Curves average 24 profiles after within-profile age-zero normalization.
Objective
Persona
PrefEval
Avg.
CE (DPPM)
54.22
86.79
70.51
+ Answer FKL
53.98
86.05
70.02
+ OPD (16)
52.96
85.80
69.38
+ OPD (64)
54.56
87.04
70.80
+ GRPO
54.80
85.06
69.93
Table 4: Training objectives on DPPM (accuracy, %). Persona denotes PersonaMem-v2. OPD (16/64) caps student rollouts at 16/64 tokens.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Optimizer
AdamW
Learning rate
10−3
Weight decay
0.01
Training epochs
5
Questions per update
1
Gradient norm clipping
1.0
Appendix
Table 5: DPPM cross-entropy training settings.
Metric
Gain (pp)
Raw p
Adj. p
PersonaMem-v2
Overall
+8.50
1.87×10−6
1.31×10−5
Self
+5.23
1.04×10−4
6.25×10−4
Current
+5.00
2.48×10−4
1.24×10−3
PrefEval
Overall
+4.32
1.50×10−3
6.02×10−3
Appendix
Table 6: DPPM versus RPMem on Qwen3-4B-Instruct-2507. Gains are in percentage points. Raw p -values use five paired training runs, and adjusted values use Holm correction across all seven metrics.
Long-running LLM agents require memory that persists and evolves across sessions. Text-based memory retrieves and reconstructs past interactions at every query, making long-horizon performance increasingly dependent on retrieval quality and contextual reasoning as histories grow. Parametric memory encodes experience directly into model computation, but existing approaches provide limited support for cross-session memory evolution. Their coupling to a specific backbone further restricts memory reuse after model replacement. We introduce RPMem, a two-stage architecture that compiles each session into a model-independent latent memory through forward computation and selectively integrates it with retained memory via a task-trained recurrent gate. The consolidated memory is then mapped to backbone-specific low-rank adaptation (LoRA) parameters, allowing the encoding capability to transfer when the backbone is replaced. Evaluation across three long-term memory benchmarks and five diverse backbones demonstrates broad generalization with near-constant update cost and memory footprint. With Qwen3-8B on PERMA, RPMem reaches 85.52%, outperforming the strongest parametric and text-based baselines by 5.32 and 12.98 percentage points, respectively. Ablations validate the complementary roles of session compilation and cross-session consolidation, while dynamics analyses reveal that the gate acquires task-specific memory integration strategies. These results establish RPMem as a lifecycle-independent parametric memory framework that maintains evolving cross-session memory that remains reusable across backbone replacements. Our implementation is available at https://github.com/Quark-Medical/rpmem/tree/main.
Personalizing large language models (LLMs) requires encoding long-term, user-specific behavioral patterns in a way that is computationally efficient, scalable, and compatible with a frozen base model. We present Latent Personal Memory (LPM), a scalable framework that represents user-specific history as a compact, persistent matrix of N latent slots, that are interpretable. A shared cross-attention projection network maps these slots into dynamic, input-conditioned soft prompts that are prepended to the input of a frozen LLM. We evaluate LPM on PersonaMem v1 and LoCOMO benchmarks across Qwen3-1.7B, 4B, and 8B backbones. Results demonstrate that LPM outperforms LoRA and Prompt Tuning by up to 8.8% and 54.4% in overall accuracy respectively on PersonaMem v1, while reducing KV-cache usage by over 64x. On LoCoMo, LPM matches LoRA accuracy with 120x fewer trainable parameters. We also show that the efficiency of LPM grows with context length and outperforms full-context at 128K context length.
Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment wastes limited memory budget on transient interactions while failing to preserve critical context for long horizon tasks. To address this gap, we investigate an underexplored question: can LLM based memory systems learn personalized memory policies? We introduce PerMemBench, the first benchmark for evaluating personalized memory systems, featuring multi year, multi domain interaction histories across diverse user personas. We further present the first empirical study of memory personalization, proposing session level storage gating, a lightweight framework that selectively bypasses memory operations for transient sessions. Our study confirms that personalization yields substantial retention gains under perfect gating, yet reveals that accurate gating remains an open and critical challenge.