Organizations: East China Normal University, Shanghai, China · Nanyang Technological University, Singapore · Zhejiang University, Hangzhou, China · Shanghai Jiao Tong University, Shanghai, China · University of Southern California, Los Angeles, USA · Northeastern University, Shenyang, China
LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions. Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set. To operationalize this view, we introduce a blinded LLM judgment over the retrieved set, together with Gold Hit and Turn Hit as evidence-coverage proxies. We then propose Budgeted Flat Reconstruction (BFR), which builds sufficient sets over a fixed flat memory store in two stages. Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them. Then, we repeatedly acquire unseen records through deeper text search or complementary entity and session views, stopping when the budget is exhausted. Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage. Specifically, on LongMemEval-S it raises judged accuracy from 72.4% to 82.2% and Turn Hit to 91.4%.
Figures & tables
Figure 1: A real error: relevant but insufficient records. The question asks how many days the remote shutter release took to arrive. Relevance ranking returns Key evidence 1 together with other photography memories, while Key evidence 2 falls below the top- k cutoff.
Figure 2: Similarity top- k vs. set-level sufficiency. (a) Prior memory architectures rank records by query similarity and return the top- k slice. (b) BFR returns a sufficient set over the flat store: it first selects complementary records that jointly cover the question’s requirements, then completes missing support under a budget by adding unseen records through text, entity, and session views.
Retrieved evidence
LLM-Suff@Set
Gold Hit
Turn Hit
Order only (Feb. 5)
0
1
0
Arrival only (Feb. 10)
0
1
1
Order + arrival
1
1
1
Table 1: Illustrative delivery-time question. Both dates are required to derive the five-day answer.
Figure 3: The two-stage evidence-construction pipeline in BFR. Stage I (FCA-MS) retrieves a candidate pool larger than the usual top- k , decomposes the query into answer requirements, prefilters candidates with FCA, and has an LLM select a compact subset that jointly covers the requirements. The selected candidates are mapped to source records to form E0 . Stage II (shown for BFR-MV instantiation) completes the set under a fixed budget: text, entity, and session access retrieve unseen records. Only new records are accumulated into Et+1 until the budget is exhausted. The comprehensive evidence set returns to Agent, and Agent answers the question based on the set.
Method
Multi-Hop
Temporal
Open Domain
Single-Hop
Overall
F1
J
F1
J
F1
J
F1
J
F1
J
Mem0 ( Chhikara et al., 2025 )
29.3
42.4
74.1
78.0
2.2
0.0
58.5
59.6
55.3
59.0
A-Mem ( Xu et al., 2025 )
32.3
42.4
70.2
72.0
3.8
0.0
47.1
51.4
48.7
53.0
CoM ∗ ( Xu et al., 2026 )
29.3
36.4
66.2
70.0
3.2
0.0
52.3
56.0
50.0
54.0
MRAgent-Flat ∗ ( Ji et al., 2026 )
32.8
45.5
72.1
78.0
2.9
0.0
63.2
68.8
58.0
64.5
BFR (ours)
44.6
48.5
78.1
80.0
2.7
12.5
75.7
77.1
68.3
70.5
Table 2: Performance across different question types on LoCoMo ( n=200 ). Same flat store and extractive answerer; BFR is BFR-MV C6. F1 and LLM-Judge (J) are percentages. Bold denotes the best value in each column.
Judge by question type
Overall
Method
Multi-S
Single-S
Temporal
Preference
Judge
Turn Hit
Mem0 ( Chhikara et al., 2025 )
60.9
87.1
65.4
26.7
72.4
64.8
A-Mem ( Xu et al., 2025 )
54.1
87.1
57.1
33.3
68.6
82.2
CoM ∗ ( Xu et al., 2026 )
54.1
87.1
57.1
33.3
68.6
80.6
MRAgent-Flat ∗ ( Ji et al., 2026 )
58.6
90.0
63.2
30.0
71.4
78.8
BFR (ours)
69.2
92.9
84.2
30.0
82.2
91.4
Table 3: Performance across different question types on LongMemEval -S ( n=500 ). The four question-type columns report Judge accuracy; All values are percentages. All rows share the same turn store, extractive answerer, and judge. Bold and underline denote the best and second-best values in each column. The full six-type Turn Hit breakdown is in Appendix K .
System
LoCoMo ( n=200 )
LongMemEval -S ( n=500 )
LLM-Suff@Set
Gold Hit
LLM-Suff@Set
Turn Hit
Mem0 ( Chhikara et al., 2025 )
45.5%
57.0%
48.0%
64.8%
A-Mem ( Xu et al., 2025 )
36.5%
50.0%
38.0%
82.2%
CoM ∗ ( Xu et al., 2026 )
44.0%
49.0%
49.0%
80.6%
MRAgent-Flat ∗ ( Ji et al., 2026 )
54.5%
62.5%
56.0%
78.8%
FCA-MS (Stage I Only)
42.5%
51.0%
48.0%
85.0%
Table 4: Set sufficiency and annotated-evidence reach across same-store systems. On LoCoMo ( n=200 ), LLM-Suff@Set is the blinded set-level judgment and Gold Hit measures annotated-evidence access. On LongMemEval -S, LLM-Suff@Set is the set-level judgment and Turn Hit measures annotated-evidence access over all 500 questions, with unannotated items counted as misses. All values are percentages. BFR-Text uses C3 and BFR-MV uses C6.
Question type
n
FCA-MS
BFR-Text
BFR-MV
CER
LLM-Suff@Set
CER
LLM-Suff@Set
CER
LLM-Suff@Set
Multi-Hop
33
6.1
15.2
18.2
27.3
12.1
30.3
Temporal
50
66.0
60.0
82.0
74.0
82.0
74.0
Open Domain
8
12.5
0.0
12.5
0.0
25.0
0.0
Single-Hop
109
41.3
45.9
71.6
71.6
70.6
73.4
Overall
200
40.5
42.5
63.0
62.0
62.0
63.5
Table 5: Where budgeted completion closes the evidence gap on LoCoMo ( n=200 ). CER is Complete Evidence Recall; LLM-Suff@Set is the blinded set-level judgment. Text and MV denote BFR-Text at C3 and BFR-MV at C6, with mean realized calls of 3.00 and 5.99 ; FCA-MS uses no completion calls. Open Domain contains only eight questions and is interpreted descriptively.
Configuration
LoCoMo
LongMemEval -S
F1
J
J
TH
FCA-MS
48.6
53.5
70.4
85.0
FCA-MS
w/ BFR-Text
67.8
70.0
81.6
88.6
FCA-MS
w/ BFR-MV
68.3
70.5
82.2
91.4
Table 6: Stage and access-schedule ablation.
Figure 4: Coverage relative to one-shot ranking.
Figure 5: Answer quality and annotated-evidence recovery versus retrieval cost on LoCoMo ( n=200 ).
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Access
Order date
Arrival date
Gold recall
Judge
Lexical top- k
✓
—
50%
Wrong
FCA-MS (Stage I)
✓
—
50%
Wrong
BFR-MV
✓
✓
100%
Correct
Appendix
Table 7: Trace behind the relevance–sufficiency case. Case evidence recall counts the two required turns; the first two sets contain relevant photography memories but miss one required date.
Multi-Hop
Temporal
Single-Hop
Overall
Method
F1
J
F1
J
F1
J
F1
J
BFR-Text †
41.4
54.5
100.0
100.0
75.5
77.3
78.7
81.5
CoM † ( Xu et al., 2026 )
41.2
63.6
96.2
96.2
71.8
84.1
75.5
85.2
MRAgent-Flat † ( Ji et al., 2026 )
12.9
45.5
92.3
92.3
62.3
68.2
65.2
72.8
Full-MRAgent † ( Ji et al., 2026 )
26.5
90.9
92.3
96.2
78.8
95.5
76.1
95.1
Appendix
Table 8: Cross-system boundary on the LoCoMo conv-30 overlap ( n=81 ), unified extractive answerer and judge; † marks rows re-scored under this protocol. Bold / underline: best / second-best. Open Domain is absent from the slice.
Configuration
Store
Gold Hit (%)
Judge Acc (%)
MRAgent-Flat
flat turns
67.9
72.8
BFR-Text
flat turns
77.8
81.5
CoM
organized pool
80.2
85.2
Full MRAgent
Cue–Tag–Content graph
88.9
95.1
Appendix
Table 9: Flat acquisition vs. structured memory on the 81 -question overlap slice under a shared extractive answerer and judge. Full MRAgent and CoM change the store and are reported as a boundary, not as fair same-substrate baselines.
Evaluation set
Parent
Size
Construction and role
LoCoMo release
Public benchmark
10 conversations
Official multi-session conversations, QA pairs, categories, and supporting turn ids.
LoCoMo pilot
LoCoMo
200 questions
Fixed stratified development subset shared by all same-store arms.
LoCoMo gold subset
LoCoMo
1,536 questions
All questions with usable gold turn ids; used for development robustness.
conv-30 overlap
LoCoMo
81 questions
Qid intersection for cross-system scoring; not a held-out split.
LLM-Suff@Set evaluations
LoCoMo pilot
1,400 arm–question pairs
200 questions paired with Mem0, A-Mem, CoM ∗ , MRAgent-Flat ∗ , FCA-MS, BFR-Text C3, and BFR-MV C6 under blinded method identities.
Budget grid
LoCoMo pilot
Fixed policy–budget runs
BFR-Text and BFR-MV call budgets evaluated before evidence-hash deduplication.
Appendix
Table 10: Provenance of the public benchmarks and derived evaluation sets. Only LoCoMo and LongMemEval-S are independently released benchmarks; all other rows are fixed subsets or evidence states constructed from them.
Method
Gold Hit
Calls
Evidence-set size
CoM ∗
49.0
1.00
10.0
MRAgent-Flat ∗
62.5
2.54
11.8
Appendix
Table 11: Retrieval cost of the adapted same-store controls on the fixed 200 -question pilot. Calls and evidence-set size are realized means.
Comparison
ΔJ
p
Δ F1
95% CI
BFR-Text − CoM ∗
+16.0
1.4×10−5
+17.8
[11.3,24.2]
BFR-MV − CoM ∗
+16.5
8.7×10−6
+18.2
[11.8,24.8]
BFR-Text − MRAgent-Flat ∗
+5.5
0.071
+9.8
[5.0,14.8]
BFR-MV − MRAgent-Flat ∗
+6.0
0.043
+10.2
[5.5,15.2]
Appendix
Table 12: Paired LoCoMo comparisons against the adapted controls. Δ is the first method minus the second. Judge uses two-sided exact McNemar; Stem F1 uses a 10,000-sample question-level paired bootstrap (seed 20260908 ).
Multi-Hop
Temporal
Open Domain
Single-Hop
Overall
Schedule
F1
J
F1
J
F1
J
F1
J
F1
J
BFR-Text
44.7
54.5
78.1
80.0
1.9
0.0
74.9
75.2
67.8
70.0
BFR-MV
44.6
48.5
78.1
80.0
2.7
12.5
75.7
77.1
68.3
70.5
Appendix
Table 13: BFR access schedules by question type on LoCoMo ( n=200 ). F1 and J are percentages; Text and MV use the C3 and C6 operating points, respectively.
FCA-MS
BFR-Text
BFR-MV
Question type
n
ER
CER
ER
CER
ER
CER
Multi-Hop
33
24.8
6.1
42.6
18.2
41.8
12.1
Temporal
50
68.0
66.0
82.0
82.0
82.0
82.0
Open Domain
8
18.8
12.5
27.1
12.5
31.2
25.0
Single-Hop
109
42.7
41.3
72.0
71.6
70.6
70.6
Overall
200
45.1
40.5
67.9
63.0
67.2
62.0
Appendix
Table 14: Evidence Recall and Complete Evidence Recall by question type on the LoCoMo pilot. Text and MV use the C3 and C6 operating points, respectively.
Type
n
FCA
Text
MV
Δ Text
Δ MV
Multi-Hop
33
15.2
27.3
30.3
+12.1
+15.2
Temporal
50
60.0
74.0
74.0
+14.0
+14.0
Open Domain
8
0.0
0.0
0.0
0.0
0.0
Single-Hop
109
45.9
71.6
73.4
+25.7
+27.5
Overall
200
42.5
62.0
63.5
+19.5
+21.0
Appendix
Table 15: Blinded LLM-Suff@Set by question type on the LoCoMo pilot. Differences use BFR minus FCA-MS. Open Domain contains only eight questions.
Method
Budget
Calls
Size
F1
J
LLM- Suff@Set
CER
FCA-MS
C0
0.00
10.00
48.6
53.5
42.5
40.5
BFR-MV
C1
1.00
21.91
65.4
67.5
57.5
59.0
C2
2.00
29.05
66.8
68.5
58.0
59.5
C3
3.00
29.84
66.8
69.5
58.0
59.5
C4
4.00
39.72
68.3
71.5
61.0
62.5
C6
5.99
39.88
68.3
70.5
63.5
62.0
Appendix
Table 16: Matched-budget quality on the LoCoMo pilot ( n=200 ). Calls and set size are realized means; size counts grounded source turns. FCA-MS selects 1 – 5 notes (mean 1.86 ) before mapping to 10 turns. J : judged accuracy; CER: Complete Evidence Recall.
Step
LoCoMo Gold Hit
LME -S Turn Hit
One-shot (FCA-MS / Mem0)
51.0
85.0 / 64.8
BFR-MV
73.5 (+22.5)
91.4 (+6.4)
BFR-Text (vs. one-shot)
72.5 (+21.5)
88.6 (+3.6)
Appendix
Table 17: Annotated-evidence reach before and after completion. LoCoMo reports Gold Hit and LongMemEval-S reports Turn Hit. The one-shot LongMemEval-S cell lists FCA-MS / Mem0; parenthesized gains use FCA-MS.
Factor
Variant
G (%)
Δ
Access
Initial
51.0
ref.
BFR-Text
72.5
+21.5
State
BFR-MV
73.5
ref.
BFR-State
73.0
−0.5
BFR-Adaptive
73.0
−0.5
Appendix
Table 18: Evidence-state ablations on the LoCoMo pilot at Bmid . Brackets: clustered 95% CI. Budgeted access drives the gain; state control does not. CI for Text vs. Initial [19.1,26.7] ; State vs. MV [−0.8,4.5] .
Method
Gold Hit (%) ↑
Initial retrieval
54.8
BFR-Text
78.0
BFR-MV
78.1
Appendix
Table 19: Development robustness on all LoCoMo questions with gold ids ( n=1536 ). Conversation-level 5 -fold; all conversations had development contact.
Method
Upd.
Mul.
Tmp.
Pref.
Asst.
User
G
J
Mem0
78.2
53.4
61.7
16.7
87.5
80.0
64.8
72.4
A-Mem
88.5
82.0
75.2
70.0
98.2
81.4
82.2
68.6
CoM
88.5
81.2
72.9
63.3
98.2
78.6
80.6
68.6
MRAgent-Flat
79.5
75.9
78.9
60.0
91.1
81.4
78.8
71.4
FCA-MS
89.7
84.2
80.5
70.0
100.0
84.3
85.0
70.4
BFR-Text
91.0
85.7
88.7
70.0
100.0
90.0
88.6
81.6
Appendix
Table 20: Full type breakdown under the unified flat-store protocol on LongMemEval -S (all 500 questions; unannotated answer-turn items count as misses). Upd.: knowledge-update; Mul.: multi-session; Tmp.: temporal; Pref.: preference; Asst./User: single-session. The six type columns and G report Turn Hit; J is judged accuracy over all questions.
Method
Gold Hit (%) ↑
Calls
Initial retrieval
51.0
0.00
BFR-Text C3
72.5
3.00
BFR-MV C3
70.5
3.00
RSC (dense residual)
58.0
3.00
Appendix
Table 21: Learned residual set completion (RSC) comparison on the LoCoMo pilot ( n=200 ). The three completion policies use a matched budget of three retrieval calls; initial retrieval is shown as a zero-call reference.
Question type
Metric
n
BFR-Text
BFR-MV
Δ
Open Domain
Judge
8
0.0
12.5
+12.5
Temporal
LLM-Suff@Set
50
74.0
70.0
−4.0
Open Domain
F1
8
1.9
3.4
+1.6
Appendix
Table 22: Selected question-type results at a matched three-call budget. LoCoMo , C3, n=200 . Values are percentages; differences are BFR-MV minus BFR-Text in percentage points, computed before rounding. Open Domain has eight questions.
Figure 6: LongMemEval -S Turn Hit under matched C3 access. BFR-MV minus BFR-Text, on 479 items with an answer-turn label; differences are percentage points.
LLM agents increasingly maintain long-term memory of user facts across sessions. Yet such memory is usually evaluated by aggregating accuracy over question rows or episodes. Because this approach scores question rows independently, even when several questions probe the same fact, it cannot show how that fact behaves as conditions change. We introduce MemTrace, a benchmark whose unit of measurement is the knowledge point: a single typed fact about the user, rather than an individual question. MemTrace probes each fact along three controlled dimensions: memory age, defined by how many sessions ago the fact appeared in the history; question type, covering current state, earlier state, and trajectory of change; and evidence condition, covering present, missing, and contradicted-by-false-premise settings. Evaluating 13 memory-system configurations across four paradigms, we find that similar pooled accuracy hides different failures: recovering a fact's current and earlier states does not imply tracking how it changed, and safe abstention does not imply correcting a false premise. The dominant bottleneck is evidence use, not retrieval: when systems fail, the evidence was retrievable 10 times more often than it was missing. These results suggest that improving long-term memory requires better use of reachable evidence, not simply more storage or retrieval.
Xianxuan Long, Zhikai Chen, Shenglai Zeng +3
Michigan State University · Case Western Reserve University
Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.
Long-term conversational large language model (LLM) agents require memory systems that can recover relevant evidence from historical interactions without overwhelming the answer stage with irrelevant context. However, existing memory systems, including hierarchical ones, still often rely solely on vector similarity for retrieval. It tends to produce bloated evidence sets: adding many superficially similar dialogue turns yields little additional recall, but lowers retrieval precision, increases answer-stage context cost, and makes retrieved memories harder to inspect and manage. To address this, we propose HiGMem (Hierarchical and LLM-Guided Memory System), a two-level event-turn memory system that allows LLMs to use event summaries as semantic anchors to predict which related turns are worth reading. This allows the model to inspect high-level event summaries first and then focus on a smaller set of potentially useful turns, providing a concise and reliable evidence set through reasoning, while avoiding the retrieval overhead that would be excessively high compared to vector retrieval. On the LoCoMo10 benchmark, HiGMem achieves the best F1 on four of five question categories and improves adversarial F1 from 0.54 to 0.78 over A-Mem, while retrieving an order of magnitude fewer turns. Code is publicly available at https://github.com/ZeroLoss-Lab/HiGMem.
Shuqi Cao, Jingyi He, Fei Tan
East China Normal University, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China