Large language model (LLM) agents that reason over clinical records must track changes in a patient's state while preserving the history needed to understand them. Simply accumulating memories leaves it unclear which information still applies, whereas overwriting earlier memories can erase evidence needed to reconstruct treatment history and clinical trajectories. We introduce STAM, a state-transition-aware memory framework that records state changes as new clinical entries arrive. STAM combines semantic retrieval with typed clinical relations to identify affected memories, maintaining current information in Active and superseded or resolved information in History. At read time, a query-dependent gate selectively serves historical memory. Across four longitudinal clinical benchmarks, we evaluate STAM with downstream question answering, direct state-maintenance diagnostics, and comparisons at approximately matched context lengths.
Figures & tables
Figure 1: Illustrative example of write-time state maintenance. In this example, Full Context and flat retrieval mix earlier positive chest-pain evidence with the current state and answer incorrectly.
Figure 2: Overview of the proposed state-aware memory framework.
Method
MMB [-1pt]mAcc.
MLC [-1pt]mF1
MQA [-1pt]F1
CB [-1pt]Acc.
Graphiti (Zep)
26.8
31.9
23.3
47.5
LightMem
35.7
33.2
22.2
50.5
A-Mem
32.5
39.0
42.3
49.2
MemOS
32.0
27.5
19.9
46.8
Mem0
22.1
28.3
23.1
45.2
MemRL
21.6
22.7
22.2
47.0
Table 1: Answer quality (%; higher is better) across persistent memory systems with Qwen3-4B-Instruct and GPT-5 readers.
Qwen
GPT-5
Data
Type
STAM
FC
STAM
FC
MMB
State update
42.0
34.0
68.0
66.0
MLC
Longitudinal
21.9
14.8
24.8
28.9
CB
Current state
56.0
50.0
66.0
62.0
Sequence
47.5
20.0
50.0
45.0
Table 2: (a) State- and time-oriented question categories. (b) Overall comparison with Full Context (FC) and Flat RAG. For the Qwen3-4B-Instruct comparisons with Flat RAG, STAM uses the approximately context-matched configurations described in Appendix C.2 .
Qwen3-4B-Instruct
GPT-5
Dataset
A+H
ACTIVE only
Δ
A+H
ACTIVE only
Δ
MMB
35.68
34.12
+1.57
51.27
46.27
+5.00
MLC
35.91
36.00
−0.09
36.73
38.28
−1.56
MQA
54.73
52.63
+2.10
85.26
53.54
+31.72
CB
50.50
49.75
+0.75
49.50
50.00
−0.50
Table 3: Historical-access ablation. Comparisons are paired within each dataset/model configuration; A+H serves both stores, whereas ACTIVE only omits History .
Method
MMB
MLC
MQA
CB
Append-only
34.16
35.40
55.99
49.25
STAM
37.40
36.20
57.45
50.20
Δ
+3.24
+0.80
+1.46
+0.95
Table 4: Paired comparison with append-only memory under the Qwen3-4B-Instruct reader using the append-only ablation configuration. Δ denotes STAM minus append-only in score points. MMB: macro accuracy; MLC: type-macro F1; MQA: Token-F1; CB: accuracy.
Variant
Pair recall ↑
False archival ↓
Full model
53.9
7.6
Semantic only, B=48
51.6
12.3
Rule: all non-episodic
28.6
17.6
Table 5: Write-time ablations: (a) MedMemoryBench and (b) MIMIC-QA. B is the impact-candidate budget.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Persons
Records
Mean tokens
Median
Range
MedMemoryBench
5
101.0
718.9K
666.4K
644.5K–815.0K
MedLoCoMo
12
31.1
37.9K
37.1K
22.6K–61.8K
MIMIC-QA (test)
20
1.1
64.7K
44.8K
9.8K–211.1K
ClinicalBench
43
5.2
20.0K
15.0K
3.4K–73.0K
Appendix
Table 6: Per-person context-length statistics for the evaluated cohorts. “Records” denotes sessions for MedMemoryBench and admissions for the remaining datasets. Token counts cover each individual’s complete record.
Benchmark
Question types
Scoring
MedMemoryBench
Entity exact match; multiple choice
Deterministic
Temporal localization; state update; inference generation; multi-hop clinical deduction
GPT-5 judge
MedLoCoMo
All six question types
Macro deterministic token-level F1
MIMIC-QA
All question types
Mean deterministic token-level F1
ClinicalBench
All nine question types
GPT-5 judge against physician-corrected references
Appendix
Table 7: Evaluation scoring by benchmark.
Benchmark
Reader
Evidence cap
Questions truncated
MedMemoryBench
GPT-5
275K
329 / 496 (66.3%)
MedLoCoMo
Qwen3-4B
100K
0 / 320 (0%)
MedLoCoMo
GPT-5
100K
0 / 320 (0%)
MIMIC-QA
Qwen3-4B
110K
179 / 400 (44.8%)
MIMIC-QA
GPT-5
280K
21 / 400 (5.2%)
ClinicalBench
Qwen3-4B
120K
0 / 400 (0%)
Appendix
Table 8: Raw-record truncation in the reported Full Context evaluations. Only final Full Context configurations reported in the main results are included.
Method
Selection
MMB
MedLoCoMo
MIMIC-QA
ClinicalBench
Qwen3-4B-Instruct
LightMem
20
1.1
0.8
1.3
1.1
A-Mem
5
23.6
8.0
18.0
8.1
MemOS
5
0.4
0.4
1.0
0.7
Mem0
5
0.1
0.1
0.1
0.1
MemRL
2 of 20
0.5
0.6
0.7
0.7
Appendix
Table 9: Baseline selection settings and median evidence tokens per question (thousands). MMB denotes MedMemoryBench. a 199 logged questions; b one persona; c Qwen3-4B memories read by GPT-5. NR: not recorded; NA: unavailable. CUPMem reports only evidence passing its premise check, so 0.0 does not imply an empty store.
Benchmark
Training
Validation
Evaluation
MedMemoryBench
Personas 6, 8–12
Persona 7
Personas 1–5
MedLoCoMo
5 patients
1 patient
12 patients; 320 questions
MIMIC-QA
9 admissions from 9 patients
1 admission from 1 patient
20 patients
ClinicalBench
30–31 patients per fold
4 patients per fold
8–9 patients per fold; 400 questions overall
Appendix
Table 10: Data partitions for write-module fine-tuning. MIMIC-QA adapter validation is an internal holdout from the benchmark’s training split. ClinicalBench combines predictions from five held-out folds.
MedMemoryBench
MedLoCoMo
MIMIC-QA
Method
Tokens/q
Recall
Tokens/q
Recall
Tokens/q
Recall
Flat RAG
1.9K
0.536
1.3K
0.560
2.3K
0.608
LightMem
1.3K
0.727
0.8K
0.624
1.3K
0.515 ‡
MemOS
0.4K
0.657
0.4K
0.543
1.0K
0.338 ‡
Mem0
0.1K
0.618
0.1K
0.543
0.1K
0.500 ‡
Ours
4.2K
0.802
4.1K
0.822
2.3K
0.959
Appendix
Table 11: Retrieval coverage under the configurations used for downstream QA. Because methods provide different amounts of context to the answering model, these results characterize their evaluation configurations rather than matched-budget retrieval comparisons. Recall uses dataset-specific gold evidence annotations and is not comparable across datasets. ClinicalBench has no evidence-level gold annotations.
Dataset
Comparator (ctx.)
STAM ctx.
Ratio
STAM
Comparator
Δ ( p )
MedMemoryBench
LightMem (1,312)
1,318
1.00
39.87
39.90
−0.03 (.95)
Flat RAG (1,898)
1,892
1.00
37.36
39.10
−1.74 (.46)
MemOS (449)
444
0.99
30.19
32.59
−2.40 (.30)
MemRL (596)
564
0.95
29.30
22.72
+6.58 (.007)
Mem0 (134)
157
1.17
24.75
22.62
+2.13 (.36)
LightMem lifted (4,505)
4,652
1.03
39.35
41.57
−2.22 (.28)
Appendix
Table 13: Context-matched evaluation. Context values are median rendered reader tokens, and Ratio is STAM context divided by comparator context. All arms are answered by the same Qwen3-4B reader under the benchmark harness’s reader settings. MedMemoryBench reports macro accuracy, MedLoCoMo reports type-macro token-F1, MIMIC-QA reports mean Token-F1, and ClinicalBench reports pooled accuracy. Rows labeled “lifted” use expanded comparator context. † Per-question comparator outputs are unavailable, so no paired significance test is reported.
Dataset
Comparator
Ratio
STAM
Cmp.
Δ ( p )
MMB
A-Mem
0.39
35.16
33.43
+1.73 (.51)
CUPMem
0.42
20.72
5.37
+15.35 ( <.001 )
MedLoCoMo
A-Mem
0.49
34.20
36.60
−2.40 (.41)
CUPMem
0.56
27.12
18.30
+8.82 ( <.001 )
MIMIC-QA
A-Mem
0.25
54.67
42.30
+12.37 ( <.001 )
CUPMem
2.12
43.96
5.92
+38.04 ( <.001 )
Appendix
Table 14: Comparisons outside the context-matching range. Ratio denotes STAM context divided by comparator context. These rows are excluded from the matched-context significance summary.
Dataset
λ
HISTORY skip
Quality Δ ( p )
MedMemoryBench
0.01
26.2%
−0.67 ( p=.289 )
MedLoCoMo
0.5
51.6%
−0.11 ( p=.469 )
MIMIC-QA
0.2
79.0%
−0.70 ( p=.147 )
ClinicalBench
0.5
1.0%
0.00 ( p=1.000 )
Appendix
Table 16: Selective HISTORY serving under the published gate configuration. The gate chooses between the original ACTIVE+HISTORY packet and the same packet with HISTORY removed. HISTORY skip is the fraction of queries served without HISTORY. Quality differences are paired against the corresponding BOTH condition.
System
Dataset
Active sufficient
History critical
History harmful
Neither correct
1.7B write / 4B reader
MMB
35.7%
4.8%
2.8%
59.5%
MLC
43.4%
0.6%
0.3%
55.9%
MQA
56.2%
4.2%
3.0%
39.5%
CB
50.2%
5.2%
5.2%
44.5%
GPT-5 write / GPT-5 reader
MMB
47.4%
11.9%
7.5%
40.7%
MLC
43.8%
4.1%
5.6%
52.2%
Appendix
Table 17: Question-level outcomes under A+H and refilled ACTIVE-only serving. A question is correct under the judge verdict for MedMemoryBench and ClinicalBench, or at token-F1 ≥0.5 for MedLoCoMo and MIMIC-QA. History harmful is a subset of Active sufficient, so the four columns do not form a disjoint partition.
Dataset
Query type
n
1.7B write / 4B reader
GPT-5
MQA
first
37
2.7%
45.9%
MQA
minimum
15
0.0%
60.0%
MQA
maximum
11
0.0%
45.5%
MQA
at_timestamp
86
4.7%
45.3%
MQA
top_n
43
0.0%
0.0%
MMB
temporal localization
100
5.0%
15.0%
Appendix
Table 18: History-critical rates for selected query types under the refilled design: the fraction of questions answered correctly with A+H but not with refilled ACTIVE-only serving.
Dataset
Retrieval
Evidence statistic
Score
MedMemoryBench
Store-aware
History/pkt 44.8%
36.3
Joint ranking
History/pkt 52.3%
35.1
MIMIC-QA
Store-aware
Gold served 97.4%
55.34
Joint ranking
Gold served 98.7%
56.23
Appendix
Table 19: Retrieval-configuration comparisons using the same underlying memories and total retrieval budget. MedMemoryBench reports macro accuracy and MIMIC-QA reports mean Token-F1. The MIMIC-QA joint-ranking condition also uses neutral evidence rendering; Table 20 separates retrieval allocation from label rendering.
Comparison
From
To
Δ
95% CI
p
Remove labels, store-aware
54.65
55.26
+0.61
[−1.8,+3.0]
.64
Add labels, joint ranking
56.80
55.54
-1.26
[−3.9,+1.3]
.35
Ranking change, both unlabelled
55.26
56.80
+1.54
[−0.4,+3.5]
.13
Ranking change, both labelled
54.65
55.54
+0.89
[−1.3,+3.1]
.43
Labelled Full → unlabelled merged
54.65
56.80
+2.15
[−0.2,+4.5]
.08
Appendix
Table 20: MIMIC-QA analysis separating retrieval allocation from store-label rendering. Scores are mean Token-F1 over 400 test questions. Δ is the paired mean difference; 95% confidence intervals are obtained by paired bootstrap and p -values by a two-sided paired sign-flip permutation test.
Variant
Archived
Pair recall
False arch.
Impact ceiling
Full, semantic+graph, B=48
34.8
53.9
7.6
71.4
Semantic only, B=48
33.0
51.6
12.3
66.7
Semantic only, B=24
25.6
42.1
17.6
52.6
Rule: measurement/dose
6.8
16.4
8.6
N/A
Rule: all non-episodic concepts
20.6
28.6
17.6
N/A
Appendix
Table 21: Additional MMB write-path controls. Impact ceiling measures coverage of the learned impact-retrieval candidate set and is therefore not applicable to deterministic supersession rules.
Criterion
Store removed
Safe recall
Safe precision
NF loss
Strict protected loss
2.3%
16.1%
92.9%
0.3%
Never-forget only
14.7%
76.3%
69.4%
8.1%
Appendix
Table 22: Effect of the calibrated loss definition on deployment-store deletion. The never-forget-only criterion is shown as a diagnostic and is not used by the system.
Candidate source
Candidates
Safe recall
Safe precision
Store removed
Updater Delete proposals
535
10.2%
84.6%
0.05%
All archived memories
8,418
16.1%
92.9%
2.3%
Appendix
Table 23: Reported deletion diagnostics by candidate source. Both runs use the same deletion gate; the updater-proposal diagnostic covers 12 personas.
Store
Macro
Pooled
EEM
IG
MCD
MQ
SUA
TLA
Untouched
36.82
39.11
63.0
22.0
8.5
41.4
44.0
42.0
After deletion
37.78
40.32
63.0
24.0
4.3
41.4
48.0
46.0
Appendix
Table 24: Paired MedMemoryBench evaluation before and after strict deletion.
Dataset
Both
Active -only
Tokens
TTFT
Tokens
TTFT
MedMemoryBench
4,482
228.7 ms
1,876
97.3 ms
MIMIC-QA
2,484
125.3 ms
1,044
59.9 ms
ClinicalBench
1,052
55.9 ms
653
41.9 ms
MedLoCoMo
1,680
90.9 ms
1,478
84.1 ms
Appendix
Table 25: Prompt length and TTFT. Values are means over the one-token prefill experiment.
Dataset
Completion tokens
Decode time
Both
Active -only
Both
Active -only
MedMemoryBench
102.8
99.5
0.877 s
0.821 s
MIMIC-QA
53.2
52.8
0.432 s
0.419 s
ClinicalBench
189.5
146.5
1.550 s
1.187 s
MedLoCoMo
8.8
8.7
0.056 s
0.054 s
Appendix
Table 26: Decode behavior on the deterministic 100-question generation subset.
Dataset
Both
Active -only
Reduction
MedMemoryBench
38.5 TFLOP
14.7 TFLOP
61.9%
MIMIC-QA
19.9 TFLOP
7.9 TFLOP
60.2%
ClinicalBench
8.0 TFLOP
4.9 TFLOP
38.9%
MedLoCoMo
13.0 TFLOP
11.4 TFLOP
12.7%
Appendix
Table 27: Prefill FLOPs of the two serving arms. These values characterize the computational cost of presenting each context to the reader and are independent of a gate’s commit frequency.
Dataset
Both
Active -only
KV reduction
Prompt-only capacity
MedMemoryBench
630 MiB
264 MiB
58%
41.6→99.3
MIMIC-QA
349 MiB
147 MiB
58%
75.0→178.4
ClinicalBench
148 MiB
92 MiB
38%
177.1→285.3
MedLoCoMo
236 MiB
208 MiB
12%
110.9→126.0
Appendix
Table 28: KV-cache footprint and prompt-only capacity under the two serving actions. Capacity is computed from the measured 186,288-token cache and ignores generated tokens and 16-token block rounding; it is therefore an upper bound on simultaneous requests, not a throughput measurement.
Dataset
Record
Source tok.
Time
MedMemoryBench
shortest
647.6K
444 s
median
668.1K
538 s
longest
818.1K
595 s
ClinicalBench
shortest
3.4K
54 s
median
15.1K
402 s
longest
73.1K
2,128 s
Appendix
Table 29: Write-time construction cost. Panel (a) reports end-to-end wall-clock latency on representative records using Qwen3-1.7B on a single otherwise idle GPU. Panel (b) reports model compute for the published Qwen3-1.7B MedMemoryBench stores over the five evaluation personas, including atomic extraction, updating, and relation linking.
1Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS) · University of Chinese Academy of Sciences, Beijing, China · 4Li Auto Inc. +1