Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density: turns are incomplete, evidence is scattered, and user states evolve over time. Existing memory methods usually rely on fixed units, such as turn-level notes or session summaries, which may lose details or introduce redundant noise. We propose FTA-Mem, a structured memory framework for low-density long-term dialogue. FTA-Mem uses Boundary-preserving Window Segmentation (BWS) to form coherent situation fragments, and constructs Fact-Time-Affect Memory Units (FTA Units) that jointly encode factual content, temporal grounding, and affective context. Retrieved units are then synthesized into structured context for answer generation. Experiments on ES-MemEval and LoCoMo show that FTA-Mem improves overall long-term memory question answering across benchmarks with different information-density characteristics. On ES-MemEval, FTA-Mem achieves 0.3871 F1 and 0.6668 BERTScore. Further analysis shows that situation-level FTA construction better balances evidence preservation and construction cost than coarse session-level or overly fine-grained turn-pair construction, providing an effective granularity trade-off for long-term dialogue memory.
Figures & tables
Figure 1: Overview of FTA-Mem. BWS segments dialogue turns into boundary-preserving fragments, which are converted into Fact-Time-Affect units. Partial units can be carried forward, adjacent units can be fused before ID assignment, and finalized units are linked, retrieved, deduplicated, and organized into structured context for grounded answer generation.
Metric
ES-MemEval
LoCoMo
Turn length
18.56
22.69
Factual anchors / turn
0.80
1.11
Entity anchors / turn
0.52
1.24
Temporal anchors / turn
0.11
0.22
Support-process anchors / turn
0.91
0.50
Implicitness (1–3)
1.91
1.58
Table 1: Turn-level information-density statistics of ES-MemEval and LoCoMo. Values are averaged over turns.
Backbone
Method
F1 Score (%) ↑
BERTScore (%) ↑
Judge ↑
Avg.R ↓
IE
TR
CD
Abs
UM
All
IE
TR
CD
Abs
UM
All
Qwen3-8B
MemoryBank
23.13
29.76
25.65
33.29
24.55
27.08
53.50
64.45
59.95
49.37
65.31
58.66
1.036
4.3
MemGPT
44.15
32.77
25.14
29.54
25.64
31.69
65.56
65.28
58.01
49.26
65.93
61.19
1.107
2.8
MemoryOS
21.41
25.79
22.35
59.96
22.29
29.70
52.35
60.83
56.73
67.48
63.87
60.09
0.917
4.9
A-Mem
27.96
31.66
26.60
39.52
25.22
29.97
56.18
65.22
60.12
53.46
65.55
60.23
1.090
3.0
CompassMem
22.06
28.85
22.06
56.52
24.32
30.20
54.09
62.35
55.69
66.98
64.70
60.67
1.024
4.4
Table 2: Performance on ES-MemEval. Avg.R is computed over the twelve automatic metrics, i.e., F1 and BERTScore on IE/TR/CD/Abs/UM/All; Judge is not included in Avg.R. Best results are in bold, and second-best results are underlined.
Backbone
Method
Single Hop
Multi Hop
Temporal
Open Domain
Average ↑
F1 ↑
BLEU-1 ↑
F1 ↑
BLEU-1 ↑
F1 ↑
BLEU-1 ↑
F1 ↑
BLEU-1 ↑
F1
BLEU-1
Qwen3-8B
MemoryBank
14.32
16.07
28.30
23.78
22.25
18.69
15.32
12.33
20.05
17.72
MemGPT
21.66
22.43
19.14
15.28
3.71
2.81
11.86
9.28
14.09
12.45
MemoryOS
18.65
22.67
26.07
22.31
3.62
2.52
14.21
11.44
15.64
14.74
A-Mem
23.92
22.69
29.42
24.85
26.81
21.85
15.68
11.90
23.96
20.32
CompassMem
31.73
29.12
42.05
35.86
34.18
28.06
18.64
14.67
31.65
26.93
Table 3: Results on LoCoMo category-level long-term memory question answering. Best results are in bold, and second-best results are underlined.
Gain
Fact
Ent.
Time
Impl.
ES-MemEval
FTA − CompassMem
-0.39
-0.12
-0.14
0.41
FTA − Avg.
-0.32
-0.07
-0.08
0.44
LoCoMo
FTA − CompassMem
0.00
0.47
-0.44
0.21
FTA − Avg.
0.26
0.40
-0.12
-0.10
Table 4: Diagnostic Pearson correlations between dialogue-level density fields and FTA-Mem’s relative F1 gain. Fact, entity, and time densities are mean annotated anchors per turn. Avg. denotes the mean over MemoryBank, MemGPT, MemoryOS, A-Mem, and CompassMem.
Table 6: Granularity and memory-unit construction cost. Units is the total number of constructed memory units. Calls and Tokens count only LLM calls and tokens used for memory-unit construction.
Figure 2: Effect of retrieval budget K on ES-MemEval and LoCoMo.
Figure 3: Ablation results on ES-MemEval and LoCoMo.
Dataset
Setting
F1 ↑
B/B1 ↑
Units
ES-MemEval
BWS
38.71
66.68
6,564
ES-MemEval
Fixed Window
37.94
65.34
6,678
LoCoMo
BWS
37.35
31.67
4,299
LoCoMo
Fixed Window
37.04
30.84
4,332
Table 7: Comparison between boundary-preserving window segmentation and fixed-window segmentation. B/B1 denotes BERTScore on ES-MemEval and BLEU-1 on LoCoMo. Scores are reported in percentages.
Setting
F1 ↑
BERTScore ↑
FTA-Mem
38.71
66.68
w/o Query Rewrite
37.44
65.90
Plain Embedding
35.99
64.88
Flat Context
37.86
65.24
w/o Auxiliary Memory
38.65
66.73
Table 8: Additional ablation results on ES-MemEval. Scores are reported in percentages.
Dataset
Gain
Field
Pearson r
Spearman ρ
95% CI
Perm. p
LOO Range
ES-MemEval
FTA–CompassMem
Fact
-0.39
-0.48
[-0.70, -0.04]
0.110
[-0.49, -0.32]
ES-MemEval
FTA–CompassMem
Implicit.
0.43
0.41
[0.00, 0.73]
0.077
[0.31, 0.53]
ES-MemEval
FTA–Avg. Baseline
Fact
-0.32
-0.36
[-0.74, 0.13]
0.154
[-0.54, -0.20]
ES-MemEval
FTA–Avg. Baseline
Implicit.
0.44
0.34
[-0.15, 0.78]
0.081
[0.23, 0.55]
LoCoMo
FTA–CompassMem
Fact
0.00
-0.20
[-0.77, 0.60]
0.481
[-0.53, 0.07]
LoCoMo
FTA–CompassMem
Implicit.
0.21
0.31
[-0.16, 0.96]
0.306
[0.26, 0.63]
Table 9: Additional density-gain correlations under the same setting as the main-paper diagnostic table. Pearson r follows the main-paper aggregation. CI denotes case-level bootstrap 95% confidence interval; Perm. p is computed by permutation testing; and LOO Range reports leave-one-case-out Pearson ranges.
Figure 4: Dialogue-level density-gain trends on ES-MemEval.
Dataset
Clusters
Mean Diff.
95% CI
ES-MemEval
18
8.43
[5.81, 11.04]
LoCoMo
10
0.55
[-0.01, 3.18]
Table 10: Paired cluster bootstrap uncertainty check for FTA-Mem minus CompassMem. Values are token-level F1 differences in percentage points. Clusters correspond to ES-MemEval cases or LoCoMo conversations. LoCoMo excludes Cat5 no-information questions.
Benchmark
Fact
Time
Affect/Ctx.
Rel.
ES-MemEval
2.00
2.00
1.48
1.36
LoCoMo
2.00
2.00
1.38
1.25
Table 11: Manual audit of intermediate FTA memory quality. Scores are reported on a 0–2 scale. Rel. denotes relation relatedness for linked memory pairs.
Query
Gold Answer
A-Mem
Ours
Q1: As of October 12, 2025 , what event made Sarah feel unsettled? ( Temporal Reasoning )
An argument with Jack about future plans .
Sarah felt unsettled due to her boss’s unexpected resignation and concerns about leadership changes.
Sarah felt unsettled after an argument with Jack about their future plans .
Q2: Has Sarah found time for painting consistently as a coping strategy? ( Conflict Detection )
No , she has found it hard to find time for painting due to work and personal issues .
Yes , Sarah has been painting consistently as a coping strategy.
Sarah finds painting calming, but struggles to find time due to work and personal issues .
Q3: What future plan indicates Jimmy’s willingness to reconnect with past relationships ? ( User Modeling )
He is considering reaching out to George for an honest conversation.
Jimmy’s willingness to reconnect is indicated by his desire to meet with Jennifer .
Jimmy plans to reach out to George for an honest conversation.
Table 12: Case study on ES-MemEval. A-Mem often retrieves related but noisy memories, while FTA-Mem better identifies the target situation, temporal state, and participant relation.
Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typically process heterogeneous dialogue content through a uniform summarization and retrieval pipeline, leading to either excessive token consumption or irreversible loss of fine-grained evidence. We argue that historical dialogue content should be handled differently according to its compressibility, temporal dynamics, and fidelity requirements. Based on this insight, we propose LeanMem, a lightweight long-term memory framework. LeanMem first filters out low-value content, then stores informative segments as compact profile memory, temporally structured event memory, or source-grounded record memory, depending on the nature of the information. During maintenance, only dynamically evolving event memories are selectively updated, avoiding redundant consolidation of stable profiles and immutable records. During inference, LeanMem dynamically selects memory types and allocates retrieval budgets according to query-specific evidence demands, assembling relevant evidence on demand. On LoCoMo and LongMemEval-S with GPT-4.1-mini and Qwen3-8B, LeanMem improves accuracy over the strongest memory-based baseline in every setting, by up to 15.1 points, at the lowest or near-lowest construction cost, inference tokens, and latency. The code and datasets are included in the supplementary materials.
Long-term memory enables LLM agents to leverage past interactions, but dialogue histories quickly exceed the context window, forcing agents to retrieve relevant subsets at query time. Because useful evidence is sparse and scattered across verbose conversations, retrieval faces a fundamental tension: broadening recall improves coverage but floods downstream reasoning with noise, while compressing memories at write time eases retrieval but irreversibly discards details that future queries may need. We introduce LazyMem, which resolves this tension by deferring all memory construction to query time. Given a retrieved candidate pool, a lightweight model processes it in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained with supervised fine-tuning followed by reinforcement learning, using a reward that jointly encourages the identification of relevant messages and the generation of compressions that are faithful to the source and useful for answering the query. On LongMemEval, LazyMem-4B achieves an LLM-judge accuracy of 0.85, outperforming the strongest non-oracle baseline while using only 213 answer-context memory tokens, 21.0 times fewer than the baseline. It further generalizes to LoCoMo without target-domain training and reduces mean latency relative to the prior query-time baseline. Code is available at https://github.com/allacnobug/LazyMem.
Jing Yu, Yibo Zhao, Jiaming Zhang +1
School of Data Science and Engineering, East China Normal University
Efficient long-term conversational memory requires retrieving sufficient evidence without indiscriminately expanding the context presented to the language model. This is challenging because relevant evidence may be distributed across multiple sessions, while compression may discard details needed for answering. Different queries therefore require different forms of memory access. To capture these demands, we formulate memory access along two dimensions: discovery breadth, which controls how broadly evidence is searched, and reading fidelity, which controls whether evidence is read in compact form or recovered from the original conversation. Based on this formulation, we introduce JustMem, which stores conversation history as compact atomic memories and adapts memory access along these two dimensions to each query. Specifically, LOOKUP handles local evidence, COMPOSE broadens discovery for distributed evidence, and REPLAY increases reading fidelity for fidelity-sensitive evidence. On LoCoMo and LongMemEval-S, JustMem achieves the highest mean accuracy and retrieval recall among the compared memory systems while using substantially fewer generative-model tokens for memory construction and inference.