Large Language Models (LLMs) serving as long-term dialogue agents require memory systems that support reliable reasoning over extended interactions. However, existing Retrieval-Augmented Generation (RAG) frameworks typically treat memory as passive storage, making it difficult to distinguish source-attributed beliefs from unattributed event/fact records and to connect evidence dispersed across sessions. We introduce CogMem, a cognitive memory architecture based on the PEC2F (Person-Event-Concept-Claim-Fact) graph schema. Dedicated Claim nodes preserve the source and target of subjective statements, while Fact and Event nodes represent semantic and episodic knowledge. Dialogue turns are incrementally converted into provenance-aware graph records, consolidated into higher-level facts, and reconciled into temporally scoped Claim views when the same source provides conflicting updates. For retrieval, a rule-based controller driven by LLM intent parsing composes four deterministic graph operators---anchoring, traversal, intersection, and evidence grounding---to reconstruct query-relevant context. Experiments on LoCoMo and LongMemEval show strong performance, especially on multi-hop, temporal, and knowledge-update tasks. Ablations and a semantic-collapse probe support complementary contributions from epistemic separation, consolidation, and agentic retrieval. Code: https://github.com/Silent-Rain02/CogMem.
Figures & tables
Figure 1: The Paradigm Shift from Passive Retrieval to Active Reconstruction. Left (traditional RAG and flat memory): passive retrieval over coarse text chunks introduces noise and may miss relevant counter-evidence. Even when a related statement is retrieved, the flat representation obscures its source and can cause the LLM to treat an attributed opinion as an unqualified record. Right (CogMem): a PEC 2 F cognitive graph supports active reconstruction, allowing the agent to attribute conflicting viewpoints to their sources and synthesize an evidence-grounded answer.
Figure 2: The Overall Architecture of CogMem. The framework consists of two layers: (1) Memory Generation & Evolution (Top) : (a) Memorize : The system processes raw dialogue to build an initial cognitive graph, converting episodic traces into graph nodes. (b) Consolidate : An offline mechanism synthesizes Facts from Events, merges duplicate Facts and Concepts, and reconciles temporally conflicting Claims into provenance-linked views without collapsing disagreements across speakers. (2) Cognitive Reasoning & Application (Bottom) : (c) Reasoning : Instead of passive retrieval, the Search Agent actively recalls information using a specialized Tool Set (Anchoring, Intersection, Traversal, Evidence Grounding) to dynamically navigate the graph and reconstruct answers through a ReAct loop.
Probe statistic
Result
Mean fact–claim cosine similarity
0.8231
Flat RAG misattribution
64/150 (42.7%)
CogMem misattribution
9/150 (6.0%)
Table 1: Direct probe of semantic collapse on 150 matched fact–claim pairs.
Comp.
Param.
Description
Value
Anchoring
α
Lexical vs. semantic balance (Eq. 2 )
0.3
Kanchor
Top-K anchor nodes
15
Traversal
Ktraverse
Max neighbors per hop
15
Dtrav
Max traversal depth
2
Intersection
Kintersect
Candidates per entity for overlap
10
Controller
Dstep
Max post-anchoring calls
5
Table 2: Hyperparameters for Cognitive Operators and Consolidation
Backbone
Method
Single Hop
Multi Hop
Temporal
Open Domain
F1
BLEU-1
F1
BLEU-1
F1
BLEU-1
F1
BLEU-1
GPT-4o-mini
Naive RAG
52.45
47.94
27.50
20.13
46.07
40.35
23.23
17.94
LightRAG
42.57
33.82
28.46
23.75
22.85
16.18
54.33
49.61
HippoRAG
39.81
31.19
39.79
37.40
26.74
22.31
51.41
50.15
RoG (Learned)
55.30
50.20
44.15
35.80
32.40
26.50
35.60
30.10
Mem0
47.65
38.72
38.72
27.13
48.93
40.51
28.64
21.58
Table 3: Performance comparison on the LoCoMo benchmark. Metrics are F1 Score and BLEU-1. Best results are in bold , and the second best are underlined .
Method
Accuracy (%)
GPT-4o-mini
Qwen2.5-14B
Naive RAG
61.00
60.80
LightRAG
52.93
47.63
HippoRAG
54.34
51.61
RoG
56.10
53.40
A-MEM
62.60
65.20
Table 4: Overall Accuracy (%) on LongMemEval . Best results are in bold , and second-best results are underlined .
Table 6: Impact of memory consolidation on accuracy (F1) and efficiency (steps).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Top-K
F1 Score (%)
Avg. Steps
K=3
52.44
2.76
K=5
54.15
2.68
K=8
54.82
2.71
K=10
55.40
2.72
K=13
56.25
2.68
K=15
56.55
2.72
Appendix
Table 7: Sensitivity analysis of the Anchoring Operator . Increasing candidate nodes ( K ) improves recall up to a saturation point, after which noise marginally degrades performance.
Max Results
F1 Score
Avg. Steps
K=5
54.96
2.72
K=10
56.31
2.72
K=15
57.04
2.69
K=20
56.90
2.67
Appendix
Table 8: Impact of the Traversal Operator’s neighbor limit (Max Results) on performance.
Max Depth
F1 Score
Avg. Steps
D=1
55.77
2.66
D=2
57.04
2.71
D=3
56.82
2.67
D=4
56.92
2.70
Appendix
Table 9: Impact of Traversal Depth . A depth of 2 provides the best trade-off for multi-hop reasoning.
Top-K
F1 Score (%)
Avg. Steps
K=3
53.78
2.68
K=5
54.07
2.72
K=8
56.44
2.73
K=10
57.04
2.72
K=13
54.67
2.68
K=15
56.15
2.68
Appendix
Table 10: Sensitivity of the Intersection Operator . Performance peaks at K=10 , indicating that commonalities are usually found in top-ranked connections.
Method
F1 (%)
Avg. Steps
CogMem (Ours)
57.0
2.65
Fixed Flow
19.7
3.00
Change
-65.4%
+13.2%
Appendix
Table 11: Ablation study of the agentic loop on the post-consolidation graph. "Fixed Flow" denotes the static 3-step retrieval pipeline.
Node Type
Before
After
Δ
Change (%)
Person
99
99
0
0.0%
Event
697
697
0
0.0%
Concept
832
819
-13
-1.6%
Fact
0
497
+497
–
Claim
2013
2069
+56
+2.8%
Total Nodes
3641
4181
+540
+14.8%
Appendix
Table 12: Graph Topology Changes . Consolidation synthesizes new Fact nodes while merging redundant Concepts.
Stage
LLM Calls
Vector Ops
Graph Ops
Indexing
O(N)
O(V)
O(V+E)
Consolidation
O(V) worst case
index updates
O(V+E) scan
Inference
O(QR)
O(QCANN(V))
O(QRdˉ)
Appendix
Table 13: Theoretical Complexity Analysis. N : dialogue turns; V,E : graph nodes and edges; Q : queries; R : controller steps; dˉ : average local degree; CANN(V) : index-dependent ANN query cost.
Method
Total Tokens (k)
LLM Calls
CogMem (Ours)
1,578.2
1,216
A-MEM
1,626.8
1,175
Mem0
1,799.4
1,614
MemoryOS
2,991.8
2,938
Appendix
Table 14: Construction Cost Comparison (LoCoMo). Total backbone tokens (input + output) used to build each memory bank with Qwen2.5-14B. CogMem uses fewer tokens than the displayed agentic baselines.
Error Source
Single-hop
Multi-hop
Temporal
Open-domain
Schema granularity (Consolidation)
8%
12%
4%
20%
Agent tool selection
24%
16%
8%
12%
Vector semantic drift
28%
20%
12%
32%
Temporal normalization
4%
8%
44%
0%
Incomplete search scope
12%
20%
12%
16%
Other/Undetermined
24%
24%
20%
20%
Appendix
Table 15: Distribution of error root causes across LoCoMo task categories (percentage of errors within each category).
Long-term conversational large language model (LLM) agents require memory systems that can recover relevant evidence from historical interactions without overwhelming the answer stage with irrelevant context. However, existing memory systems, including hierarchical ones, still often rely solely on vector similarity for retrieval. It tends to produce bloated evidence sets: adding many superficially similar dialogue turns yields little additional recall, but lowers retrieval precision, increases answer-stage context cost, and makes retrieved memories harder to inspect and manage. To address this, we propose HiGMem (Hierarchical and LLM-Guided Memory System), a two-level event-turn memory system that allows LLMs to use event summaries as semantic anchors to predict which related turns are worth reading. This allows the model to inspect high-level event summaries first and then focus on a smaller set of potentially useful turns, providing a concise and reliable evidence set through reasoning, while avoiding the retrieval overhead that would be excessively high compared to vector retrieval. On the LoCoMo10 benchmark, HiGMem achieves the best F1 on four of five question categories and improves adversarial F1 from 0.54 to 0.78 over A-Mem, while retrieving an order of magnitude fewer turns. Code is publicly available at https://github.com/ZeroLoss-Lab/HiGMem.
Shuqi Cao, Jingyi He, Fei Tan
East China Normal University, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China
Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.
Large Language Models (LLMs) have made significant progress in dialogue, yet redundant memory contexts severely limit their effectiveness in long-term dialogue agents. External memory systems have been proposed to improve memory maintenance. However, these systems mainly rely on one-shot retrieval, which limits their ability to retrieve sufficient and relevant evidence. Although recent methods introduce reflection into retrieval, their retrieval paths are generated by the LLM from limited evidence, leading to unstable retrieval and additional latency overhead. %These limitations highlight the need for effective retrieval mechanisms. To address these limitations, we propose MGRetrieval, a retrieval strategy that grounds reflective retrieval in the semantic structure of historical memories. Specifically, MGRetrieval consists of two steps: (1) It references the structure of historical memories to construct a more precise retrieval path. (2) The LLM retains critical memories and determines whether accumulated memories are sufficient to stop further iterative retrieval. This allows the retrieval process to follow semantically meaningful paths. Through memory-guided retrieval and critical memory propagation, MGRetrieval gradually constructs concise and sufficient memory contexts. Extensive experiments on LoCoMo show that MGRetrieval outperforms the strongest baseline by 8.91% in F1 and 11.11% in BLEU-1 on average across Qwen2.5-14B and Qwen3-14B, while maintaining practical token and latency costs. The code can be found in https://anonymous.4open.science/r/MGRetrieval.
Tan Wang, Yunwei Dong
Northwestern Polytechnical University Xi’an, Shaanxi, China