LLM-based agents face challenges in personalized interactive tasks due to heterogeneous, multi-typed, and implicitly constrained long-term traces. Existing memory mechanisms struggle with accurate routing and retrieval, especially on-device where personalization is critical. Most methods use single-vector representations, blurring type distinctions and relational structure. We propose HGP, a hybrid graph memory framework. HGP employs a lightweight self-enhancement classifier for personalized memory routing and constructs episodic, semantic, and procedural memories as graphs. It also extracts working memory as a state trajectory to capture current state and implicit constraints, ensuring reliable decision-making. The classifier reduces large-model calls, enabling on-device deployment, while graph storage enables accurate retrieval and incremental user profile refinement. Experiments on two benchmarks show that on PAL-Set solution selection, HGP achieves an S-score of 35.58, nearly 7 points above the strongest baseline. Code and data are at https://github.com/Ouan6/HGP-.git.
Figures & tables
Figure 1: Overview of our HGP framework with two components: (1) a self-enhancement classifier that routes inputs via a classifier/LLM and incrementally updates the classifier using high-confidence samples, and (2) hybrid graph storage that preserves structural properties of heterogeneous memory types.
Group
Method
B-1
B-2
B-3
B-4
QA Judge
S-score
Δ (vs. Vanilla w/o log)
Non-memory
Vanilla (w/o log)
19.68
7.72
3.80
2.11
7.31
20.34
+0.00
Vanilla (with log)
18.96
7.51
3.82
2.22
7.32
25.00
+4.66
Turn-level RAG
19.48
7.72
3.93
2.23
7.41
25.06
+4.72
Session-level RAG
19.17
7.55
3.82
2.22
7.30
25.24
+4.90
Memory-based
MemoryBank
20.83
7.36
4.23
2.56
7.13
28.85
+8.51
Mem0
19.20
7.66
3.82
2.12
7.34
26.50
+6.16
Table 2: Performance comparison across different methods on PAL-Set Solution QA tasks. B-1 to B-4 represent BLEU scores. Δ (vs. Vanilla w/o log) reports the absolute improvement in S-score compared to the Vanilla (w/o log) baseline. LLM-judge denotes the score evaluated by GPT-4o-mini.
Figure 2: Working Memory Analysis: working-memory state tracking in a personal healthcare scenario. Yellow highlights denote pending to-do reminders, and blue highlights indicate preference-aligned personalized suggestions. Compared with baselines, our method surfaces the working-memory state (pre-meal medication: pending) and reminds the user while providing personalized dinner recommendations.
Method
B-1
B-2
B-3
B-4
LLM-judge
S-score
Unified-vector
18.15
7.24
3.60
2.05
7.18
25.15
Typed-vector
18.61
7.51
3.74
2.09
7.29
33.66
w/o E-Mem
19.53
7.96
3.96
2.23
7.09
28.83
w/o S-Mem
19.69
8.30
4.38
2.57
7.23
25.71
w/o P-Mem
19.59
8.63
4.66
2.76
7.31
33.10
Full (Ours)
19.62
8.67
4.70
2.78
7.46
35.58
Table 3: Ablation study on memory components and representation designs. The table reports generation quality (BLEU and LLM-judge) and selection performance (S-score). Higher is better.
Figure 3: Classifier routing statistics across iterations. (a) Routing mix and adaptive threshold. (b) Cumulative workload handled by the LLM and the classifier.
Router
Micro-F1
Macro-F1
Avg Latency (ms)
Offline Classifier
0.8638
0.8790
4.06
Fixed LLM Router
0.8953
0.9224
1803.69
Zero-shot LLM Router
0.8286
0.8436
1572.64
Ours
0.9105
0.9339
4.98
Table 4: Performance comparison of the self-enhancement classifier against other baselines. Best results are in bold , and second-best results are underlined .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: LLM-as-a-Judge prompt used for PAL-Set Solution QA evaluation, enforcing requirement-fit scoring with implicit constraints, retrieved memories/logs, and pos/neg reference calibration.
Rank
τ0
β
τmin
F1 (mean ± std)
Latency (ms)
Norm Latency
Score
1
0.9
0.02
0.5
0.9490±0.0034
4.23
1.000
0.947
2
0.9
0.02
0.3
0.9490±0.0034
4.36
1.031
0.932
3
0.9
0.02
0.7
0.9490±0.0034
4.40
1.040
0.927
4
0.8
0.1
0.7
0.9482±0.0023
4.46
1.055
0.919
5
0.9
0.05
0.5
0.9474±0.0060
4.44
1.050
0.919
6
0.8
0.02
0.5
0.9482±0.0023
4.72
1.118
0.888
Appendix
Table 5: Hyperparameter sensitivity analysis of the router with different values of τ0 , β , and τmin . The configuration selected for our experiments is highlighted in bold.
Figure 5: An illustrative example of HGP in a personalized health management scenario. Given a user’s current query and long-term interaction history, HGP organizes heterogeneous memories, including episodic, semantic, procedural, and working memory, within a hybrid graph storage, and retrieves task-relevant evidence to generate personalized health recommendations.
Stage
Macro-F1
Micro-F1
Δ Macro (vs. S0)
S0 (offline init)
0.8980
0.8859
0.0000
S1
0.9162
0.9068
+0.0182
S2
0.9282
0.9192
+0.0302
S3
0.9445
0.9275
+0.0465
S4
0.9527
0.9379
+0.0547
S5
0.9554
0.9416
+0.0574
Appendix
Table 6: Performance evolution of the self-enhancement classifier across incremental training stages.
Figure 6: Special case study: illustrating the detailed process by which an LLM retrieves user-relevant memories via HGP in the special case. Working memory addresses implicit constraints in a personalized health assistant scenario; its effect is highlighted in yellow (pre-meal medication: pending), while personalized recommendation cues are highlighted in blue.
Figure 7: Illustration of the procedural abstraction operator Asem .
Figure 8: Illustration of the procedural abstraction operator Aproc .