The memory capabilities of Large Language Models (LLMs) have garnered increasing attention recently. Despite great success achieved, existing retrieval-based memory approaches typically overlook the differences between memories and employ a unified strategy to process all memories, leading to suboptimal performance. Thus, an intuitive question arises: can we categorize memory into different types and select appropriate strategies? However, given the topic-rich, scenario-complex, and boundary-blurred nature of memory scenarios, achieving precise classification of memories is not easy. To address this challenge, we propose a memory multi-class dataset in this paper, termed TriMEM, which provides precise annotations for memory types across diverse scenarios. Building upon this foundation, we propose a novel memory framework, named MemoType, which can adaptively recognize each memory and query type with the learned router model. With the memory and query routing, MemoType can retrieve the memory with corresponding query types and design tailored retrieval strategies, thereby enhancing the retrieval performance. Moreover, we theoretically prove that any single retrieval strategy is subject to a fundamental upper bound on its expected retrieval precision in multi-class corpora, leading to systematic precision degradation. Extensive experiments on three datasets demonstrate that MemoType consistently outperforms existing methods, achieving up to 16.18% improvement in Recall@1.
Figures & tables
Figure 1 : Evidence and Motivation: Semantic Clustering of memory and the MemoType framework.
Figure 2 : Three Memory Types in TriMEM (Left) and the proposed MemoType Framework (Right).
Method
Recall@1
NDCG@1
Recall@3
NDCG@3
Recall@5
NDCG@5
Recall@10
NDCG@10
LongMemEval-S
HyDE ( Gao et al., 2023 )
27.23
27.23
53.19
37.82
65.32
42.84
82.34
49.83
Mill ( Jia et al., 2024 )
38.72
38.72
67.66
50.17
78.94
56.10
90.64
61.34
Query2Doc ( Wang et al., 2023a )
25.96
25.96
50.64
36.43
64.89
41.66
81.06
48.38
SeCom ( Pan et al., 2025 )
45.32
45.32
74.04
53.42
83.19
59.42
91.28
64.66
HippoRAG2 ( Gutiérrez et al., 2025 )
50.64
50.64
80.85
60.59
88.51
66.48
93.53
71.05
Table 1 : Retrieval Performance. All methods are based on Contriever as the retriever.
Method
GPT4Judge
F1
BLEU
Rouge1
Rouge2
RougeL
RougeLsum
BERTScore
LongMemEval-S
HyDE ( Gao et al., 2023 )
42.60
9.96
1.44
10.58
4.67
9.31
9.57
83.02
Mill ( Jia et al., 2024 )
42.20
9.91
1.55
10.57
4.54
9.20
9.45
82.96
Query2Doc ( Wang et al., 2023a )
43.00
9.80
1.47
10.44
4.46
9.21
9.43
82.97
SeCom ( Pan et al., 2025 )
44.80
11.01
1.65
11.66
5.14
10.36
10.58
83.26
HippoRAG2 ( Gutiérrez et al., 2025 )
45.40
10.55
1.65
11.15
5.17
9.98
10.17
83.12
Table 2 : Question Answer Performance. All methods are based on Contriever as the retriever.
Strategy
LongMemEval-S
LoCoMo
PerLTQA
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Key Expansion
51.77
76.62
85.59
25.93
43.25
51.06
66.10
83.64
88.05
Hypothetical Query
53.91
78.29
87.06
24.42
43.00
50.10
62.65
82.90
89.08
Hypothetical Memory
52.03
77.45
86.22
25.23
42.30
51.11
66.24
83.88
88.34
Type-Aware Strategy
59.15
83.19
90.21
31.07
47.73
54.08
71.92
88.05
92.78
Improvement
5.24
4.90
3.15
5.14
4.48
2.97
5.68
4.17
3.70
Table 3 : The ablation study of the proposed type-aware strategy with single retrieval strategy on three datasets.
Figure 3 : The ablation study of the memory pruning module on two benchmarks.
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
Memory Type
Content Type
Example
EM
Event Occurred
“Sarah and John had dinner at a sushi restaurant downtown.”
Events Planned to Occur
“Next Month, my colleagues are planning to see a beautiful sunset.”
PM
User Information
“Jack’s degree is in Mathematics. ”
User Preferences and Habits
“I love Italian food, especially pasta and pizza.”
User Intent
“I want to buy a pair of running shoes”
GM
General Knowledge
“Water boils at 100 degrees Celsius at sea level.”
Appendix
Table 4 : The Memory Classification Criterion in the TriMEM Benchmark. “EM” denotes Episodic Memory, “PM” denotes the Personal Semantic Memory, “GM” denotes the General Semantic Memory.
Figure 4 : The Statistical Information of TriMEM 3 3 3 The percentage of the memory types is a relative ratio, as there is overlap between different memory types. .
Figure 9
Figure 7 : The Contriever Similarity Distribution between TriMem Context with the raw context on two benchmarks.
Type
Model
TriMEM(In Distribution)
PerLTQA(Out of Distribution)
Overall
EM
PM
GM
Profile
Relationship
Events
ACC (100 %)
Prompting
Qwen3-8b
71.90
81.65
71.37
62.68
88.65
57.21
46.23
Prompting
Gemini-2.5-Flash
77.84
82.47
76.72
74.33
100.00
75.06
91.83
Prompting
GPT-4o-Mini
83.55
85.05
87.18
78.42
100.00
89.47
92.00
Prompting
Qwen3-32b
77.75
87.82
80.88
64.55
100.00
95.64
87.55
Appendix
Table 5 : The classification Performance on TriMEM with four LLM models. “EM” denotes Episodic Memory, “PM” denotes the Personal Semantic Memory, “GM” denotes the General Semantic Memory.
Figure 8 : The Distribution of three Similarity Metrics (Word Overlap Rate, BLEU Score, and 3-gram Overlap Rate) between Transferred Memory and Original Memory on LongMemEval-S and LoCoMo datsets. The red dotted line indicates a similarity threshold of 0.5. The majority of samples fall in the low similarity region, indicating a significant difference between the transformed text and the original text.
Figure 9 : The Frequency Distributions of three Similarity Metrics between Transferred Memory and Original Memory on LoCoMo dataset. The red/green dashed lines indicate the mean and median, respectively. The distribution is overall low and long-tailed, further indicating that the transformed memory differs significantly from the original.
Figure 10 : The Frequency Distributions of three Similarity Metrics between Transferred Memory and Original Memory on LongMemEval-S dataset. The red/green dashed lines indicate the mean and median, respectively. It is evident that all three indicators are concentrated in the lower range.
Dataset
#Inst.
#QA
#Sess.
Sess./Inst.
Turns/Sess.
Words/Inst.
#Q-Types
LoCoMo
10
1,986
272
27.2
21.6
15.9 k
5
LongMemEval-S
500
500
25,097
50.2
0 9.8
79.0 k
6
LongMemEval-M
500
500
249,525
499.1
0 9.8
777.6 k
6
PerLTQA
32
2,838
128
4.0
21.7
0 8.9 k
4
Appendix
Table 6 : Statistics of the long-term memory QA datasets used in our study. “#Inst.” counts the top-level units. “Sess./Inst.” and “Turns/Sess.” are averages; “Words/Inst.” is the average number of whitespace-tokenized words per instance.
Method
LongMemEval-S
LoCoMo
PerLTQA
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Base Retriever: MPNet
HyDE
25.11
53.19
66.60
27.29
44.76
53.52
30.62
65.68
74.17
Mill
27.45
53.19
64.89
21.00
36.81
44.51
36.15
67.27
74.77
Query2Doc
27.45
54.68
68.72
26.38
46.12
54.03
33.51
65.82
74.63
SeCom
37.66
62.13
73.19
21.45
34.24
41.49
36.19
65.72
73.75
Appendix
Table 7 : The ablation study of different retrievers.
Method
Recall@1
NDCG@1
Recall@3
NDCG@3
Recall@5
NDCG@5
Recall@10
NDCG@10
LongMemEval-S
w/o Classify
43.40
43.40
73.62
54.00
83.83
60.33
94.26
65.79
w/ Classify
50.21
50.21
79.79
60.32
86.81
65.47
92.34
70.15
LoCoMo
w/o Classify
20.75
20.75
36.56
30.86
44.06
34.16
53.93
37.31
w/ Classify
25.13
25.13
41.99
35.88
49.70
39.06
59.47
42.36
Appendix
Table 8 : The ablation study of the proposed classification strategy on four datasets. “w/ Classify” represents adopting the classification strategy to filter memory, “w/o Classify” denotes the opposite.
Strategy
GPT4Judge
F1
BLEU
Rouge1
Rouge2
RougeL
RougeLsum
BERTScore
LongMemEval-S
w/o Prun
55.60
19.55
4.48
21.32
10.17
19.54
19.70
85.03
w/ Prun
56.60
20.31
4.91
22.18
10.70
20.28
20.32
85.22
LoCoMo
w/o Prun
48.39
20.30
4.11
21.36
10.63
19.93
19.91
85.55
w/ Prun
49.19
22.23
4.64
23.28
11.83
21.83
21.86
85.80
Appendix
Table 9 : The ablation study of memory pruning strategy on three datasets. “w/ Prun” denotes adopting the proposed memory pruning strategy, and“w/o Prun” denotes the contrary.
Figure 11 : Strategy Ablation on Retrieval Recall@1 across Three Benchmarks.
Figure 12 : Strategy Ablation on Retrieval Recall@3 across Three Benchmarks.
Method
Recall@1
NDCG@1
Recall@3
NDCG@3
Recall@5
NDCG@5
Recall@10
NDCG@10
LongMemEval-S
BERT
55.74
55.74
81.06
63.40
88.72
68.64
93.62
72.54
DeBERTa
55.32
55.32
81.91
63.19
89.36
68.90
95.74
72.62
RoBERTa
58.09
58.09
83.83
65.27
90.21
70.70
95.53
74.41
Qwen3-1.7B
59.15
59.15
83.19
66.74
90.21
71.60
92.77
74.85
LoCoMo
Appendix
Table 10 : The Ablation Study of BERT Encoder.
Dataset
Method
Calls
Calls/sample
Prompt (M)
Compl. (M)
Total (M)
USD
LoCoMo
A-Mem
11,764
1,176.4
5.73
1.82
7.56
$1.95
HippoRAG2
11,760
1,176.0
3.97
0.33
4.30
$0.79
MemoType
1,756
175.6
0.29
0.08
0.37
$0.09
LongMemEval-S
A-Mem
493,574
987.1
326.67
76.50
403.18
$94.90
HippoRAG2
492,146
984.3
254.34
29.21
283.56
$55.68
MemoType
19,897
39.8
3.78
1.43
5.21
$1.42
Appendix
Table 11 : The cost analysis across four benchmarks. “Calls” is the number of paid LLM calls during memory ingestion. Due to the time limits, † HippoRAG2 on LongMemEval-M is linearly projected from 5/500 samples.
Figure 13 : The Semantic Clustering Evidence with Contriever Embedding on two Benchmarks.
Figure 14 : Transfer Prompt for User-Assistant Conversation
Figure 15 : Transfer Prompt for User-User Conversation
Figure 16 : Episodic Memory Classification Prompt
Figure 17 : Personal Semantic Memory Classification Prompt
Figure 18 : General Semantic Memory Classification Prompt
Figure 19 : Memory Multi-Classification Prompt
Figure 20 : HyDE Prompt
Figure 21 : MILL Prompt
Figure 22 : Query2Doc Prompt
Figure 23 : Answer Generation Prompt for User-Assistant Conversation
Figure 24 : Answer Generation Prompt for User-User Conversation
As Large Language Models (LLMs) are increasingly used for long-duration tasks, maintaining effective long-term memory has become a critical challenge. Current methods often face a trade-off between cost and accuracy. Simple storage methods often fail to retrieve relevant information, while complex indexing methods (such as memory graphs) require heavy computation and can cause information loss. Furthermore, relying on the working LLM to process all memories is computationally expensive and slow. To address these limitations, we propose MemSifter, a novel framework that offloads the memory retrieval process to a small-scale proxy model. Instead of increasing the burden on the primary working LLM, MemSifter uses a smaller model to reason about the task before retrieving the necessary information. This approach requires no heavy computation during the indexing phase and adds minimal overhead during inference. To optimize the proxy model, we introduce a memory-specific Reinforcement Learning (RL) training paradigm. We design a task-outcome-oriented reward based on the working LLM's actual performance in completing the task. The reward measures the actual contribution of retrieved memories by mutiple interactions with the working LLM, and discriminates retrieved rankings by stepped decreasing contributions. Additionally, we employ training techniques such as Curriculum Learning and Model Merging to improve performance. We evaluated MemSifter on eight LLM memory benchmarks, including Deep Research tasks. The results demonstrate that our method meets or exceeds the performance of existing state-of-the-art approaches in both retrieval accuracy and final task completion. MemSifter offers an efficient and scalable solution for long-term LLM memory. We have open-sourced the model weights, code, and training data to support further research.
Jiejun Tan, Zhicheng Dou, Liancheng Zhang +3
Gaoling School of Artificial Intelligence · Renmin University of China · Beijing, China
Long-term memory enables LLM agents to leverage past interactions, but dialogue histories quickly exceed the context window, forcing agents to retrieve relevant subsets at query time. Because useful evidence is sparse and scattered across verbose conversations, retrieval faces a fundamental tension: broadening recall improves coverage but floods downstream reasoning with noise, while compressing memories at write time eases retrieval but irreversibly discards details that future queries may need. We introduce LazyMem, which resolves this tension by deferring all memory construction to query time. Given a retrieved candidate pool, a lightweight model processes it in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained with supervised fine-tuning followed by reinforcement learning, using a reward that jointly encourages the identification of relevant messages and the generation of compressions that are faithful to the source and useful for answering the query. On LongMemEval, LazyMem-4B achieves an LLM-judge accuracy of 0.85, outperforming the strongest non-oracle baseline while using only 213 answer-context memory tokens, 21.0 times fewer than the baseline. It further generalizes to LoCoMo without target-domain training and reduces mean latency relative to the prior query-time baseline. Code is available at https://github.com/allacnobug/LazyMem.
Jing Yu, Yibo Zhao, Jiaming Zhang +1
School of Data Science and Engineering, East China Normal University
Long-term conversational large language model (LLM) agents require memory systems that can recover relevant evidence from historical interactions without overwhelming the answer stage with irrelevant context. However, existing memory systems, including hierarchical ones, still often rely solely on vector similarity for retrieval. It tends to produce bloated evidence sets: adding many superficially similar dialogue turns yields little additional recall, but lowers retrieval precision, increases answer-stage context cost, and makes retrieved memories harder to inspect and manage. To address this, we propose HiGMem (Hierarchical and LLM-Guided Memory System), a two-level event-turn memory system that allows LLMs to use event summaries as semantic anchors to predict which related turns are worth reading. This allows the model to inspect high-level event summaries first and then focus on a smaller set of potentially useful turns, providing a concise and reliable evidence set through reasoning, while avoiding the retrieval overhead that would be excessively high compared to vector retrieval. On the LoCoMo10 benchmark, HiGMem achieves the best F1 on four of five question categories and improves adversarial F1 from 0.54 to 0.78 over A-Mem, while retrieving an order of magnitude fewer turns. Code is publicly available at https://github.com/ZeroLoss-Lab/HiGMem.
Shuqi Cao, Jingyi He, Fei Tan
East China Normal University, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China