The memory capabilities of Large Language Models (LLMs) have garnered increasing attention recently. Despite great success achieved, existing retrieval-based memory approaches typically overlook the differences between memories and employ a unified strategy to process all memories, leading to suboptimal performance. Thus, an intuitive question arises: can we categorize memory into different types and select appropriate strategies? However, given the topic-rich, scenario-complex, and boundary-blurred nature of memory scenarios, achieving precise classification of memories is not easy. To address this challenge, we propose a memory multi-class dataset in this paper, termed TriMEM, which provides precise annotations for memory types across diverse scenarios. Building upon this foundation, we propose a novel memory framework, named MemoType, which can adaptively recognize each memory and query type with the learned router model. With the memory and query routing, MemoType can retrieve the memory with corresponding query types and design tailored retrieval strategies, thereby enhancing the retrieval performance. Moreover, we theoretically prove that any single retrieval strategy is subject to a fundamental upper bound on its expected retrieval precision in multi-class corpora, leading to systematic precision degradation. Extensive experiments on three datasets demonstrate that MemoType consistently outperforms existing methods, achieving up to 16.18% improvement in Recall@1.
Figures & tables
Figure 1 : Evidence and Motivation: Semantic Clustering of memory and the MemoType framework.
Figure 2 : Three Memory Types in TriMEM (Left) and the proposed MemoType Framework (Right).
Method
Recall@1
NDCG@1
Recall@3
NDCG@3
Recall@5
NDCG@5
Recall@10
NDCG@10
LongMemEval-S
HyDE ( Gao et al., 2023 )
27.23
27.23
53.19
37.82
65.32
42.84
82.34
49.83
Mill ( Jia et al., 2024 )
38.72
38.72
67.66
50.17
78.94
56.10
90.64
61.34
Query2Doc ( Wang et al., 2023a )
25.96
25.96
50.64
36.43
64.89
41.66
81.06
48.38
SeCom ( Pan et al., 2025 )
45.32
45.32
74.04
53.42
83.19
59.42
91.28
64.66
HippoRAG2 ( Gutiérrez et al., 2025 )
50.64
50.64
80.85
60.59
88.51
66.48
93.53
71.05
Table 1 : Retrieval Performance. All methods are based on Contriever as the retriever.
Method
GPT4Judge
F1
BLEU
Rouge1
Rouge2
RougeL
RougeLsum
BERTScore
LongMemEval-S
HyDE ( Gao et al., 2023 )
42.60
9.96
1.44
10.58
4.67
9.31
9.57
83.02
Mill ( Jia et al., 2024 )
42.20
9.91
1.55
10.57
4.54
9.20
9.45
82.96
Query2Doc ( Wang et al., 2023a )
43.00
9.80
1.47
10.44
4.46
9.21
9.43
82.97
SeCom ( Pan et al., 2025 )
44.80
11.01
1.65
11.66
5.14
10.36
10.58
83.26
HippoRAG2 ( Gutiérrez et al., 2025 )
45.40
10.55
1.65
11.15
5.17
9.98
10.17
83.12
Table 2 : Question Answer Performance. All methods are based on Contriever as the retriever.
Strategy
LongMemEval-S
LoCoMo
PerLTQA
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Key Expansion
51.77
76.62
85.59
25.93
43.25
51.06
66.10
83.64
88.05
Hypothetical Query
53.91
78.29
87.06
24.42
43.00
50.10
62.65
82.90
89.08
Hypothetical Memory
52.03
77.45
86.22
25.23
42.30
51.11
66.24
83.88
88.34
Type-Aware Strategy
59.15
83.19
90.21
31.07
47.73
54.08
71.92
88.05
92.78
Improvement
5.24
4.90
3.15
5.14
4.48
2.97
5.68
4.17
3.70
Table 3 : The ablation study of the proposed type-aware strategy with single retrieval strategy on three datasets.
Figure 3 : The ablation study of the memory pruning module on two benchmarks.
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
Memory Type
Content Type
Example
EM
Event Occurred
“Sarah and John had dinner at a sushi restaurant downtown.”
Events Planned to Occur
“Next Month, my colleagues are planning to see a beautiful sunset.”
PM
User Information
“Jack’s degree is in Mathematics. ”
User Preferences and Habits
“I love Italian food, especially pasta and pizza.”
User Intent
“I want to buy a pair of running shoes”
GM
General Knowledge
“Water boils at 100 degrees Celsius at sea level.”
Appendix
Table 4 : The Memory Classification Criterion in the TriMEM Benchmark. “EM” denotes Episodic Memory, “PM” denotes the Personal Semantic Memory, “GM” denotes the General Semantic Memory.
Figure 4 : The Statistical Information of TriMEM 3 3 3 The percentage of the memory types is a relative ratio, as there is overlap between different memory types. .
Figure 9
Figure 7 : The Contriever Similarity Distribution between TriMem Context with the raw context on two benchmarks.
Type
Model
TriMEM(In Distribution)
PerLTQA(Out of Distribution)
Overall
EM
PM
GM
Profile
Relationship
Events
ACC (100 %)
Prompting
Qwen3-8b
71.90
81.65
71.37
62.68
88.65
57.21
46.23
Prompting
Gemini-2.5-Flash
77.84
82.47
76.72
74.33
100.00
75.06
91.83
Prompting
GPT-4o-Mini
83.55
85.05
87.18
78.42
100.00
89.47
92.00
Prompting
Qwen3-32b
77.75
87.82
80.88
64.55
100.00
95.64
87.55
Appendix
Table 5 : The classification Performance on TriMEM with four LLM models. “EM” denotes Episodic Memory, “PM” denotes the Personal Semantic Memory, “GM” denotes the General Semantic Memory.
Figure 8 : The Distribution of three Similarity Metrics (Word Overlap Rate, BLEU Score, and 3-gram Overlap Rate) between Transferred Memory and Original Memory on LongMemEval-S and LoCoMo datsets. The red dotted line indicates a similarity threshold of 0.5. The majority of samples fall in the low similarity region, indicating a significant difference between the transformed text and the original text.
Figure 9 : The Frequency Distributions of three Similarity Metrics between Transferred Memory and Original Memory on LoCoMo dataset. The red/green dashed lines indicate the mean and median, respectively. The distribution is overall low and long-tailed, further indicating that the transformed memory differs significantly from the original.
Figure 10 : The Frequency Distributions of three Similarity Metrics between Transferred Memory and Original Memory on LongMemEval-S dataset. The red/green dashed lines indicate the mean and median, respectively. It is evident that all three indicators are concentrated in the lower range.
Dataset
#Inst.
#QA
#Sess.
Sess./Inst.
Turns/Sess.
Words/Inst.
#Q-Types
LoCoMo
10
1,986
272
27.2
21.6
15.9 k
5
LongMemEval-S
500
500
25,097
50.2
0 9.8
79.0 k
6
LongMemEval-M
500
500
249,525
499.1
0 9.8
777.6 k
6
PerLTQA
32
2,838
128
4.0
21.7
0 8.9 k
4
Appendix
Table 6 : Statistics of the long-term memory QA datasets used in our study. “#Inst.” counts the top-level units. “Sess./Inst.” and “Turns/Sess.” are averages; “Words/Inst.” is the average number of whitespace-tokenized words per instance.
Method
LongMemEval-S
LoCoMo
PerLTQA
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Recall@1
Recall@3
Recall@5
Base Retriever: MPNet
HyDE
25.11
53.19
66.60
27.29
44.76
53.52
30.62
65.68
74.17
Mill
27.45
53.19
64.89
21.00
36.81
44.51
36.15
67.27
74.77
Query2Doc
27.45
54.68
68.72
26.38
46.12
54.03
33.51
65.82
74.63
SeCom
37.66
62.13
73.19
21.45
34.24
41.49
36.19
65.72
73.75
Appendix
Table 7 : The ablation study of different retrievers.
Method
Recall@1
NDCG@1
Recall@3
NDCG@3
Recall@5
NDCG@5
Recall@10
NDCG@10
LongMemEval-S
w/o Classify
43.40
43.40
73.62
54.00
83.83
60.33
94.26
65.79
w/ Classify
50.21
50.21
79.79
60.32
86.81
65.47
92.34
70.15
LoCoMo
w/o Classify
20.75
20.75
36.56
30.86
44.06
34.16
53.93
37.31
w/ Classify
25.13
25.13
41.99
35.88
49.70
39.06
59.47
42.36
Appendix
Table 8 : The ablation study of the proposed classification strategy on four datasets. “w/ Classify” represents adopting the classification strategy to filter memory, “w/o Classify” denotes the opposite.
Strategy
GPT4Judge
F1
BLEU
Rouge1
Rouge2
RougeL
RougeLsum
BERTScore
LongMemEval-S
w/o Prun
55.60
19.55
4.48
21.32
10.17
19.54
19.70
85.03
w/ Prun
56.60
20.31
4.91
22.18
10.70
20.28
20.32
85.22
LoCoMo
w/o Prun
48.39
20.30
4.11
21.36
10.63
19.93
19.91
85.55
w/ Prun
49.19
22.23
4.64
23.28
11.83
21.83
21.86
85.80
Appendix
Table 9 : The ablation study of memory pruning strategy on three datasets. “w/ Prun” denotes adopting the proposed memory pruning strategy, and“w/o Prun” denotes the contrary.
Figure 11 : Strategy Ablation on Retrieval Recall@1 across Three Benchmarks.
Figure 12 : Strategy Ablation on Retrieval Recall@3 across Three Benchmarks.
Method
Recall@1
NDCG@1
Recall@3
NDCG@3
Recall@5
NDCG@5
Recall@10
NDCG@10
LongMemEval-S
BERT
55.74
55.74
81.06
63.40
88.72
68.64
93.62
72.54
DeBERTa
55.32
55.32
81.91
63.19
89.36
68.90
95.74
72.62
RoBERTa
58.09
58.09
83.83
65.27
90.21
70.70
95.53
74.41
Qwen3-1.7B
59.15
59.15
83.19
66.74
90.21
71.60
92.77
74.85
LoCoMo
Appendix
Table 10 : The Ablation Study of BERT Encoder.
Dataset
Method
Calls
Calls/sample
Prompt (M)
Compl. (M)
Total (M)
USD
LoCoMo
A-Mem
11,764
1,176.4
5.73
1.82
7.56
$1.95
HippoRAG2
11,760
1,176.0
3.97
0.33
4.30
$0.79
MemoType
1,756
175.6
0.29
0.08
0.37
$0.09
LongMemEval-S
A-Mem
493,574
987.1
326.67
76.50
403.18
$94.90
HippoRAG2
492,146
984.3
254.34
29.21
283.56
$55.68
MemoType
19,897
39.8
3.78
1.43
5.21
$1.42
Appendix
Table 11 : The cost analysis across four benchmarks. “Calls” is the number of paid LLM calls during memory ingestion. Due to the time limits, † HippoRAG2 on LongMemEval-M is linearly projected from 5/500 samples.
Figure 13 : The Semantic Clustering Evidence with Contriever Embedding on two Benchmarks.
Figure 14 : Transfer Prompt for User-Assistant Conversation
Figure 15 : Transfer Prompt for User-User Conversation
Figure 16 : Episodic Memory Classification Prompt
Figure 17 : Personal Semantic Memory Classification Prompt
Figure 18 : General Semantic Memory Classification Prompt
Figure 19 : Memory Multi-Classification Prompt
Figure 20 : HyDE Prompt
Figure 21 : MILL Prompt
Figure 22 : Query2Doc Prompt
Figure 23 : Answer Generation Prompt for User-Assistant Conversation
Figure 24 : Answer Generation Prompt for User-User Conversation