Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.
Figures & tables
Figure 1: Conversational vs. organizational memory. Top: a single narrator restates their own facts, and LLM extraction at ingest keeps only the surviving fact. Bottom: different authors in an organization write the same fact, and an earlier decision may still hold.
Figure 2: Overview of Mem++ framework. (3.2) Each document is stored whole with its author, date and sentence embedding, and it is indexed three ways. (3.3) Records valid at the as-of date θ are retrieved from the three indexes and fused with weighted RRF, and three of the k slots are reserved for the latest-dated matches. (3.4) The top- k records are passed with their dates and authors to a fixed answering model. Mem g ++ graph triples carry no date.
Figure 3: The six capabilities of OrgMemBench ( Gardner, 2026 ) , with questions and answers examples. The tier we evaluated holds 443 artifacts across 157 threads for 18 months, with 73 questions.
Method
Supersession
Decision Provenance
Bi-temporal
Audit Replay
Justification Chain
Contradiction
Overall
gpt-4.1-mini
Full Context
0.0
27.5
0.0
33.0
22.7
8.3
17.8 ± 0.1
RAG
66.1
79.7
66.7
40.9
19.6
75.4
55.0 ± 0.3
Zep
31.4
42.5
23.0
76.8
20.7
43.3
41.0 ± 0.2
Mem0
40.8
55.0
50.0
25.1
14.4
52.8
36.9 ± 0.0
A-Mem
58.3
63.3
43.6
36.5
20.7
43.5
44.5 ± 0.3
gbrain
54.8
60.7
30.3
48.5
22.7
22.9
43.2 ± 0.1
Table 1: Performance on OrgMemBench ( Gardner, 2026 ) categorized by question type. Overall is weighted by the number of questions in each category and reported with its standard deviation across runs. Bold indicates the best performance. Underline indicates the second best performance.
Method
Temporal Reasoning
Open Domain
Multi-Hop
Single-Hop
Average
↑ LLM
↑ F1
↑ BLEU
↑ LLM
↑ F1
↑ BLEU
↑ LLM
↑ F1
↑ BLEU
↑ LLM
↑ F1
↑ BLEU
↑ LLM
↑ F1
↑ BLEU
gpt-4.1-mini
Full Context
74.2
47.5
40.0
56.6
28.4
22.2
77.2
44.2
33.7
86.9
61.4
53.4
80.6
53.3
45.0
RAG-2048
66.8
50.3
40.1
48.6
28.3
21.7
67.4
39.1
28.9
82.8
57.7
48.1
74.5
50.9
41.3
RAG-4096
27.4
22.3
19.1
28.8
17.9
13.9
31.7
20.1
12.8
35.9
25.8
22.0
32.9
23.5
19.2
Zep
60.2
23.9
20.0
43.8
24.2
19.3
53.7
30.5
20.4
66.9
45.5
40.0
61.6
36.9
30.9
Mem0
56.9
39.2
33.2
47.9
23.7
17.7
68.2
40.1
30.3
71.4
48.6
42.0
66.3
43.5
36.5
Table 2: Performance on LoCoMo ( Maharana et al., 2024 ) categorized by type. Bold, underline indicates best and second best performance. The baseline performance comes from Nan et al. (2025) .
Question Type
Full Context
Zep
Nemori
Mem++
Mem g ++
gpt-4o-mini
Single-session preference
6.7
20.0
46.7
46.7
50.0
Single-session assistant
89.3
80.4
83.9
94.6
96.4
Temporal reasoning
42.1
62.4
61.7
56.7
56.7
Multi-session
38.3
57.9
51.1
63.4
62.0
Knowledge update
78.2
83.3
61.5
84.3
85.2
Single-session user
78.6
92.9
88.6
98.4
98.4
Table 3: Performance on LongMemEval S ( Wu et al., 2024 ) by question type. LLM-judge accuracy is reported. Bold indicates the best performance. Underline indicates the second best performance.
Variant
OrgMemBench
LoCoMo
LongMemEval S
Mem++
57.0
85.7
87.9
w/o vector leg
21.3 ( − 35.7)
76.7 ( − 9.1)
11.9 ( − 76.0)
w/o lexical leg
57.2 ( + 0.2)
85.2 ( − 0.5)
88.7 ( + 0.7)
w/o tag leg
55.0 ( − 2.0)
–
–
w/ consolidation
55.6 ( − 1.3)
85.4 ( − 0.3)
86.9 ( − 1.1)
w/ fact index
55.2 ( − 1.7)
83.1 ( − 2.6)
–
Table 4: Ablation with claude-sonnet-4-6 as the answerer. Parentheses give the change from Mem++.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
OrgMemBench (medium)
LoCoMo
LongMemEval S
Category
Count
Category
Count
Category
Count
C1 Supersession
15
Single-hop
841
Knowledge update
72
C2 Provenance
15
Multi-hop
282
Multi-session
121
C3 Bi-temporal
7
Temporal
321
Temporal reasoning
127
C4 Audit replay
15
Open domain
96
Single-session user
64
C5 Justification
15
–
–
Single-session assistant
56
Appendix
Table 5: Category distributions of the OrgMemBench ( Gardner, 2026 ) , LoCoMo ( Maharana et al., 2024 ) and LongMemEval S ( Wu et al., 2024 ) datasets.
Setting
OrgMemBench
LoCoMo
LongMemEval S
Embedding model
all-MiniLM-L6-v2
text-embedding-3-small
text-embedding-3-small
Embedding size
384
1536
1536
Weights (wlex,wtag,wvec)
(1, 1, 4)
(1, 1, 2)
(1, 1, 4)
Candidates per index C
50
200
50
Lexical matching
conjunctive
disjunctive
conjunctive
Appendix
Table 6: Retrieval settings that differ across benchmarks.
Figure 4: OrgMemBench score of Mem++ as the number of retrieved rows k varies. The dashed line is Full Context, and the dotted line marks k=50 used in the main experiments.
k
gpt-4o-mini
gpt-4.1-mini
Retrieved tokens
5
36.57
43.99
0.20M
10
45.10
50.56
0.41M
15
47.16
54.44
0.61M
30
45.94
56.49
1.22M
50
44.23
57.00
2.01M
70
48.67
56.70
2.77M
Appendix
Table 7: OrgMemBench score of Mem++ for each k . Retrieved tokens are summed over all questions and counted before the 40,000-character cut.