Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to accumulate history and to serve readout, so what the memory stores cannot be controlled separately from what it exposes to the current computation. We propose LSTMem, an LSTM-inspired online memory that instead equips each layer of a frozen LLM with two matrix-valued states: a cell state that accumulates history and a hidden state whose readouts correct the backbone's attention. Input and forget gates control what the cell stores, while an output gate separately controls what the cell exposes through the hidden state. LSTMem further connects memory across depth through forward hidden-state propagation and block-end feedback, and uses higher-layer reconstruction gradients to refine lower-layer cell states before rebuilding hidden states from shallow to deep layers. Across memory benchmarks on Qwen3-4B-Instruct, LSTMem consistently improves MemoryAgentBench, LoCoMo, and HotpotQA over the plain backbone. Comparisons further show that the LSTM-based memory formulation outperforms an associative-memory counterpart, while removing cross-layer hidden-memory propagation degrades performance. These results demonstrate the benefits of separating memory accumulation from memory expression and organizing memory hierarchically across model depth. The code is available at https://github.com/Longchentong/LSTMem.
Figures & tables
Model
HotpotQA ↑
MemoryAgentBench ↑
EM
F1
Avg.
AR
TTL
LRU
SF
Qwen3-4B-Instruct
42.35
56.00
29.54
35.30
26.14
47.08
14.37
External Memory
+ BM25 RAG
40.35−2.00
52.83−3.17
24.43−5.11
32.20−3.10
9.74−16.40
37.86−9.22
15.00+0.63
+ LLMLingua-2
36.93−5.42
50.03−5.97
15.63−13.91
21.45−13.85
1.43−24.71
38.45−8.63
8.62−5.75
+ MemoryBank
—
—
17.65−11.89
22.65−12.65
7.67−18.47
36.36−10.72
9.88−4.49
Table 1: Main results across three base models. Bold marks the highest score within each model group; subscripts give point changes from its unadapted model. Baseline scores are from Lei et al. (2026b) . HotpotQA is scored by exact match (EM) and token-level F1 against the gold answer. AR, TTL, LRU and SF denote accurate retrieval, test-time learning, long-range understanding and selective forgetting. AR and SF are scored by accuracy, TTL by classification accuracy and movie-recommendation Recall@5, and LRU by summarization F1 and Detective QA accuracy. MAB averages use 2,000/700/171/800 question weights for AR/TTL/LRU/SF; baseline averages are approximated from rounded family scores.
Model
Avg.
Multi
Temp.
Open
Single
Qwen3-4B-Instruct
40.79
38.39
32.89
10.77
48.05
+ BM25 RAG
36.68
38.12
20.34
9.99
45.47
+ LLMLingua-2
40.98
39.07
30.13
10.98
49.19
+ MemoryBank
38.14
37.88
21.76
13.35
47.31
+ Context2LoRA
48.11
37.95
34.99
16.75
60.11
+ MemGen
40.05
32.93
33.30
12.67
48.13
Table 2: LoCoMo F1 (%). Avg. is weighted over 1,540 questions. Multi, Temp., Open and Single denote multi-hop, temporal, open-domain and single-hop questions.
Model
Avg.
Multi
Temp.
Open
Single
Qwen3-4B-Instruct
40.79
38.39
32.89
10.77
48.05
+ BM25 RAG
36.68
38.12
20.34
9.99
45.47
+ LLMLingua-2
40.98
39.07
30.13
10.98
49.19
+ MemoryBank
38.14
37.88
21.76
13.35
47.31
+ Context2LoRA
48.11
37.95
34.99
16.75
60.11
+ MemGen
40.05
32.93
33.30
12.67
48.13
Table 2: LoCoMo F1 (%). Avg. is weighted over 1,540 questions. Multi, Temp., Open and Single denote multi-hop, temporal, open-domain and single-hop questions.
Table 4: Depth ablation. Δ = on − off. The coupling and feedback are removed in training and inference alike.
Seed
Depth coupling
MAB Avg.
AR
TTL
LRU
SF
42
on
45.08
55.70
47.50
37.59
18.00
off
43.16
54.30
41.29
38.21
18.00
Δ
+1.92
+1.40
+6.21
−0.62
0.00
43
on
45.45
54.90
52.36
37.44
17.50
off
43.68
54.35
44.36
37.39
17.75
Δ
+1.77
+0.55
+8.00
+0.05
−0.25
Table 4: Depth ablation. Δ = on − off. The coupling and feedback are removed in training and inference alike.
Memory layout
Avg.
Multi
Temp.
Open
Single
All 36
53.60
46.25
40.94
19.92
64.75
Even 18
51.30
45.97
40.51
16.10
61.22
Odd 18
51.56
46.83
39.73
14.10
61.93
First 12
48.09
43.91
33.98
11.12
59.10
Middle 12
50.42
46.93
39.19
13.94
60.03
Last 12
50.58
47.50
41.37
14.28
59.28
Table 5: Memory placement ablation on LoCoMo. F1 (%) on Qwen3-4B-Instruct. Layer indices are one-based: even/odd select 18 alternating layers; first/middle/last select layers 1–12, 13–24 and 25–36, respectively.
Method
Training stage
Multi-hop
Temporal
Open-domain
Single-hop
Avg.
δ -Mem
1: QASPER
42.57
39.31
18.12
58.59
49.12
2: + Long
46.73
38.75
16.20
59.36
50.06
LSTMem
1: QASPER
43.70
39.03
18.72
60.19
50.18
2: + Long
46.25
40.94
19.92
64.75
53.60
Table 6: Training-stage comparison on LoCoMo. Stage 1 trains on QASPER, and Stage 2 continues training on the Long dataset. Scores are F1 (%), with Avg. weighted by question count.
Block-end feedback
Avg.
Multi
Temp.
Open
Single
On
53.60
46.25
40.94
19.92
64.75
Off
51.89
46.31
40.78
17.48
61.93
Table 7: Block-end feedback ablation on LoCoMo. F1 (%) on Qwen3-4B-Instruct. Both rows keep forward depth coupling and differ only in whether block-end feedback is applied.