Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized graph configurations traverse a common base graph; AdaptiveGraph adds chronological edges and Personalized PageRank diffusion. We also evaluate BM25 over the same extracted notes and OpenClaw as a raw-input external reference. Retrieval rankings vary across memory settings. On LongMemEval-S, AdaptiveGraph is the strongest graph configuration at 0.844 MRR, but BM25 reaches 0.867 and OpenClaw 0.880. On ATANT Core, localized graph traversal outperforms diffusion and BM25, whereas BM25 leads the stress rounds. Reducing LongMemEval-S within the tested range does not reproduce the ATANT diffusion penalty, but the smallest tested store remains larger than ATANT Core, so store size cannot be ruled out. The penalty also persists under a permissive content-match criterion. Vocabulary normalization and extraction quality substantially affect graph retrieval, and missing extraction tags are common among top-five misses. Retrieval strategies should therefore be evaluated jointly with the memory setting and against strong lexical baselines.
Figures & tables
Benchmark
Extraction
Basic
Stop
TFIDF
Adapt
BM25
OpenClaw (orig.)
LongMemEval-S
GPT-OSS
0.296
0.404
0.428
0.768
0.794
0.880
LongMemEval-S
Qwen3-Small
0.263
0.419
0.435
0.817
0.847
0.880
LongMemEval-S
Gemma
0.495
0.706
0.718
0.844
0.867
0.880
LongMemEval (oracle)
all three
1.000
0.996–0.998
1.000
1.000
1.000
0.990
ATANT (Core)
GPT-OSS
0.682
0.703
0.691
0.544
0.641
0.705
ATANT (Core)
Qwen3-Small
0.729
0.734
0.708
0.554
0.631
0.705
Table 1: Retrieval MRR across benchmarks and extraction models. Basic, Stop, and TFIDF are localized TagGraph configurations; Adapt is AdaptiveGraph. BM25 ranks the shared extracted notes without a graph, while OpenClaw indexes raw conversations as an external reference. Bold marks the best method on the shared extracted store. Stress is the mean over R2–R5.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Extraction
Cases
Notes/case
Distinct tags/case
Tags/note
Reuse
Gemma
496
249.5
376.5
4.95
3.27
GPT-OSS
469
232.1
918.5
6.50
1.64
Qwen3-Small
482
240.1
1,140.5
9.29
1.96
Appendix
Table 2: Tag vocabulary statistics on LongMemEval-S. Reuse is tag occurrences divided by distinct tags within a case, averaged over cases: a higher value means the same tags are used again instead of new ones being minted.
Dataset
− Diffusion
− Chronological
− File seeding
LongMemEval-S
−0.020
−0.001
−0.119
ATANT (Core)
+0.158
+0.077
−0.009
Appendix
Table 3: Change in MRR from removing one AdaptiveGraph component (GPT-OSS extraction). Negative means the component was contributing; positive means removing it improves accuracy.
Strict
Permissive
Extraction
Full
− Diffusion
Penalty
Full
− Diffusion
Penalty
GPT-OSS
0.546
0.705
+0.158
0.570
0.713
+0.143
Qwen3-Small
0.547
0.699
+0.153
0.574
0.714
+0.140
Gemma
0.549
0.699
+0.150
0.577
0.712
+0.135
Appendix
Table 4: ATANT Core MRR with and without diffusion under both content-match criteria ( n=304 questions). Penalty is the MRR gain from removing diffusion, computed before rounding; MRR values use the scoring implementation of Section 5.1 and agree with Table 1 within 0.007.
Retrieval
MRR
Token F1
Containment
TagGraph-Stopwords
0.409
0.058
14.6%
TagGraph-TFIDF
0.463
0.072
18.7%
AdaptiveGraph
0.773
0.086
25.3%
Appendix
Table 5: Downstream answer accuracy on LongMemEval-S (GPT-OSS extraction, Qwen3-Small generation, N=487 ). Containment: the normalized gold answer appears inside the generated answer.
Question type
Extraction
k
Basic
Stop
TFIDF
Adapt
OC-notes
single-session-user
GPT-OSS
1
24.3
30.0
34.3
88.6
82.9
3
45.7
55.7
58.6
98.6
97.1
5
55.7
65.7
67.1
98.6
97.1
Qwen3-Small
1
20.0
31.4
27.1
87.1
82.9
3
37.1
60.0
58.6
97.1
94.3
5
47.1
67.1
67.1
97.1
94.3
Appendix
Table 6
Category
n
B
S
T
A
Career
32
0.779
0.844
0.823
0.608
Daily Life
41
0.739
0.773
0.827
0.632
Health
31
0.685
0.702
0.677
0.640
Learning
36
0.753
0.749
0.731
0.540
Life Events
107
0.639
0.649
0.646
0.489
Relationships
57
0.620
0.644
0.583
0.499
Appendix
Table 7: ATANT Core MRR by category (GPT-OSS extraction). B/S/T = TagGraph-Basic/-Stopwords/-TFIDF, A = AdaptiveGraph.
GPT-OSS
Qwen3-Small
Gemma
Split
Local
Adapt
Local
Adapt
Local
Adapt
Core
0.703
0.544
0.734
0.554
0.695
0.554
R2
0.764
0.757
0.764
0.752
0.779
0.770
R3
0.825
0.805
0.810
0.781
0.811
0.807
R4
0.861
0.851
0.848
0.839
0.841
0.853
R5
0.857
0.872
0.869
0.863
0.864
0.872
Appendix
Table 8: ATANT MRR by split, best localized variant versus AdaptiveGraph.
Diffusion penalty
Split
Questions
Notes/story
GPT-OSS
Qwen3-Small
Gemma
Core
304
7.7
+0.158
+0.153
+0.150
R2
367
6.7
+0.075
+0.080
+0.073
R3
386
6.2
+0.085
+0.094
+0.079
R4
380
5.1
+0.049
+0.049
+0.049
R5
398
4.9
+0.035
+0.026
+0.029
Appendix
Table 9: ATANT questions, store size, and diffusion penalty by split. Penalty is the strict-MRR gain from removing diffusion (iteration count zero), computed before rounding.
Category
nLongMem
nATANT
Gemma
GPT-OSS
Qwen3-Small
C1
398
212
0.0 / 0.0
0.0 / 0.0
0.0 / 0.0
C2
0
41
43.9 / 100.0
36.6 / 100.0
53.7 / 100.0
C3
0
0
—
—
—
C4
0
7
100.0 / 85.7
100.0 / 100.0
100.0 / 100.0
Appendix
Table 10: Failure diagnostics over the 658-case pooled failure subset. Cases are included if at least one of the four graph retrieval configurations under at least one extraction model misses the gold note at recall@5. C1–C4 are defined in the text; examples appear in Appendix J . Model cells give Tagged / Succeeded percentages within the category.
Institute for Artificial Intelligence, Peking University · School of Intelligence Science and Technology, Peking University · School of Computer Science and Technology, Beijing Institute of Technology