Large Language Models (LLMs) serving as long-term dialogue agents require memory systems that support reliable reasoning over extended interactions. However, existing Retrieval-Augmented Generation (RAG) frameworks typically treat memory as passive storage, making it difficult to distinguish source-attributed beliefs from unattributed event/fact records and to connect evidence dispersed across sessions. We introduce CogMem, a cognitive memory architecture based on the PEC2F (Person-Event-Concept-Claim-Fact) graph schema. Dedicated Claim nodes preserve the source and target of subjective statements, while Fact and Event nodes represent semantic and episodic knowledge. Dialogue turns are incrementally converted into provenance-aware graph records, consolidated into higher-level facts, and reconciled into temporally scoped Claim views when the same source provides conflicting updates. For retrieval, a rule-based controller driven by LLM intent parsing composes four deterministic graph operators---anchoring, traversal, intersection, and evidence grounding---to reconstruct query-relevant context. Experiments on LoCoMo and LongMemEval show strong performance, especially on multi-hop, temporal, and knowledge-update tasks. Ablations and a semantic-collapse probe support complementary contributions from epistemic separation, consolidation, and agentic retrieval. Code: https://github.com/Silent-Rain02/CogMem.
Figures & tables
Figure 1: The Paradigm Shift from Passive Retrieval to Active Reconstruction. Left (traditional RAG and flat memory): passive retrieval over coarse text chunks introduces noise and may miss relevant counter-evidence. Even when a related statement is retrieved, the flat representation obscures its source and can cause the LLM to treat an attributed opinion as an unqualified record. Right (CogMem): a PEC 2 F cognitive graph supports active reconstruction, allowing the agent to attribute conflicting viewpoints to their sources and synthesize an evidence-grounded answer.
Figure 2: The Overall Architecture of CogMem. The framework consists of two layers: (1) Memory Generation & Evolution (Top) : (a) Memorize : The system processes raw dialogue to build an initial cognitive graph, converting episodic traces into graph nodes. (b) Consolidate : An offline mechanism synthesizes Facts from Events, merges duplicate Facts and Concepts, and reconciles temporally conflicting Claims into provenance-linked views without collapsing disagreements across speakers. (2) Cognitive Reasoning & Application (Bottom) : (c) Reasoning : Instead of passive retrieval, the Search Agent actively recalls information using a specialized Tool Set (Anchoring, Intersection, Traversal, Evidence Grounding) to dynamically navigate the graph and reconstruct answers through a ReAct loop.
Probe statistic
Result
Mean fact–claim cosine similarity
0.8231
Flat RAG misattribution
64/150 (42.7%)
CogMem misattribution
9/150 (6.0%)
Table 1: Direct probe of semantic collapse on 150 matched fact–claim pairs.
Comp.
Param.
Description
Value
Anchoring
α
Lexical vs. semantic balance (Eq. 2 )
0.3
Kanchor
Top-K anchor nodes
15
Traversal
Ktraverse
Max neighbors per hop
15
Dtrav
Max traversal depth
2
Intersection
Kintersect
Candidates per entity for overlap
10
Controller
Dstep
Max post-anchoring calls
5
Table 2: Hyperparameters for Cognitive Operators and Consolidation
Backbone
Method
Single Hop
Multi Hop
Temporal
Open Domain
F1
BLEU-1
F1
BLEU-1
F1
BLEU-1
F1
BLEU-1
GPT-4o-mini
Naive RAG
52.45
47.94
27.50
20.13
46.07
40.35
23.23
17.94
LightRAG
42.57
33.82
28.46
23.75
22.85
16.18
54.33
49.61
HippoRAG
39.81
31.19
39.79
37.40
26.74
22.31
51.41
50.15
RoG (Learned)
55.30
50.20
44.15
35.80
32.40
26.50
35.60
30.10
Mem0
47.65
38.72
38.72
27.13
48.93
40.51
28.64
21.58
Table 3: Performance comparison on the LoCoMo benchmark. Metrics are F1 Score and BLEU-1. Best results are in bold , and the second best are underlined .
Method
Accuracy (%)
GPT-4o-mini
Qwen2.5-14B
Naive RAG
61.00
60.80
LightRAG
52.93
47.63
HippoRAG
54.34
51.61
RoG
56.10
53.40
A-MEM
62.60
65.20
Table 4: Overall Accuracy (%) on LongMemEval . Best results are in bold , and second-best results are underlined .
Table 6: Impact of memory consolidation on accuracy (F1) and efficiency (steps).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Top-K
F1 Score (%)
Avg. Steps
K=3
52.44
2.76
K=5
54.15
2.68
K=8
54.82
2.71
K=10
55.40
2.72
K=13
56.25
2.68
K=15
56.55
2.72
Appendix
Table 7: Sensitivity analysis of the Anchoring Operator . Increasing candidate nodes ( K ) improves recall up to a saturation point, after which noise marginally degrades performance.
Max Results
F1 Score
Avg. Steps
K=5
54.96
2.72
K=10
56.31
2.72
K=15
57.04
2.69
K=20
56.90
2.67
Appendix
Table 8: Impact of the Traversal Operator’s neighbor limit (Max Results) on performance.
Max Depth
F1 Score
Avg. Steps
D=1
55.77
2.66
D=2
57.04
2.71
D=3
56.82
2.67
D=4
56.92
2.70
Appendix
Table 9: Impact of Traversal Depth . A depth of 2 provides the best trade-off for multi-hop reasoning.
Top-K
F1 Score (%)
Avg. Steps
K=3
53.78
2.68
K=5
54.07
2.72
K=8
56.44
2.73
K=10
57.04
2.72
K=13
54.67
2.68
K=15
56.15
2.68
Appendix
Table 10: Sensitivity of the Intersection Operator . Performance peaks at K=10 , indicating that commonalities are usually found in top-ranked connections.
Method
F1 (%)
Avg. Steps
CogMem (Ours)
57.0
2.65
Fixed Flow
19.7
3.00
Change
-65.4%
+13.2%
Appendix
Table 11: Ablation study of the agentic loop on the post-consolidation graph. "Fixed Flow" denotes the static 3-step retrieval pipeline.
Node Type
Before
After
Δ
Change (%)
Person
99
99
0
0.0%
Event
697
697
0
0.0%
Concept
832
819
-13
-1.6%
Fact
0
497
+497
–
Claim
2013
2069
+56
+2.8%
Total Nodes
3641
4181
+540
+14.8%
Appendix
Table 12: Graph Topology Changes . Consolidation synthesizes new Fact nodes while merging redundant Concepts.
Stage
LLM Calls
Vector Ops
Graph Ops
Indexing
O(N)
O(V)
O(V+E)
Consolidation
O(V) worst case
index updates
O(V+E) scan
Inference
O(QR)
O(QCANN(V))
O(QRdˉ)
Appendix
Table 13: Theoretical Complexity Analysis. N : dialogue turns; V,E : graph nodes and edges; Q : queries; R : controller steps; dˉ : average local degree; CANN(V) : index-dependent ANN query cost.
Method
Total Tokens (k)
LLM Calls
CogMem (Ours)
1,578.2
1,216
A-MEM
1,626.8
1,175
Mem0
1,799.4
1,614
MemoryOS
2,991.8
2,938
Appendix
Table 14: Construction Cost Comparison (LoCoMo). Total backbone tokens (input + output) used to build each memory bank with Qwen2.5-14B. CogMem uses fewer tokens than the displayed agentic baselines.
Error Source
Single-hop
Multi-hop
Temporal
Open-domain
Schema granularity (Consolidation)
8%
12%
4%
20%
Agent tool selection
24%
16%
8%
12%
Vector semantic drift
28%
20%
12%
32%
Temporal normalization
4%
8%
44%
0%
Incomplete search scope
12%
20%
12%
16%
Other/Undetermined
24%
24%
20%
20%
Appendix
Table 15: Distribution of error root causes across LoCoMo task categories (percentage of errors within each category).