Organizations: South China University of Technology · University of the Chinese Academy of Sciences · Beijing University of Posts and Telecommunications · University of Science and Technology of China · Tsinghua University · Peking University · Fudan University · Shanghai Jiao Tong University
Long-term agents face growing storage demands as they accumulate experience. World models capture reusable regularities that can reduce the information stored for each experience. We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content. Its task-aware allocation rule balances the expected impact of reconstruction errors against storage cost, retaining information with downstream value beyond the predictive prior. Across five long-term agent-memory benchmarks, MemoWM achieves 42.42% average answer accuracy, exceeding the strongest baseline by 2.62 percentage points, while reducing average experience-specific storage by 53.9% relative to MIRIX, the most storage-efficient baseline. Further analysis shows that stronger world models reduce per-experience storage at comparable task quality. Accounting for model parameters reveals a trade-off between shared model capacity and recurring storage costs, with the capacity that minimizes total storage increasing as more interactions are retained. Our code is available at https://github.com/Feld-maxiu/MemoWM.
Figures & tables
Figure 1: Average accuracy versus experience-specific storage.
Figure 2: Overview of the MemoWM framework. Coordinates are retained when their estimated task impact exceeds the weighted storage cost; others are filled by a causal world model. (A) Task-relevant retention, (B) shared-history reconstruction, and (C) offline fitting and calibration. Reconstructed states are retrieved through the Reader Bridge to answer future queries.
Method
WorldMemArena
AMA-Bench
LME-V2
WebQuest
MemEye
Avg.
RAG
52.97 ± 0.12
44.87 ± 0.70
40.25 ± 0.69
22.35 ± 0.50
35.42 ± 0.93
39.17
UniversalRAG
40.38 ± 0.19
40.35 ± 0.88
34.56 ± 1.18
21.68 ± 0.75
37.95 ± 1.24
34.98
MEM1
45.66 ± 0.31
14.96 ± 1.64
29.41 ± 1.83
20.25 ± 1.01
32.05 ± 1.93
28.47
MemAgent
43.03 ± 0.37
29.18 ± 1.60
29.33 ± 1.91
16.90 ± 0.78
30.88 ± 2.34
29.86
A-Mem
55.37 ± 0.28
34.85 ± 1.31
41.28 ± 1.48
26.70 ± 0.85
36.33 ± 1.28
38.91
MemGPT
58.43 ± 0.20
35.44 ± 0.81
44.97 ± 1.29
24.55 ± 0.87
35.61 ± 1.16
39.80
Table 1: Main results across five long-horizon memory benchmarks. (a) The upper table reports answer accuracy (%); (b) the middle table reports the experience-specific storage rate (kbit/observation); (c) the break-even point is the number of retained observations at which MemoWM’s total storage is no greater than the listed comparator’s. LME-V2 denotes LongMemEval-V2, Omni-SM denotes Omni-SimpleMem, and Avg. is the unweighted mean across benchmarks.
Figure 3: (a) Upper left: all-coordinate coding rate and (b) upper right: gated total rate with amortized FP16 weights, both versus model size at three training-data fractions. (c) Lower left: all-coordinate rate under four predictive contexts; blue curves show increases over full context (upper axis). (d) Lower right: storage-optimal gated model size versus retained transitions.
Figure 4: Left: where decoder-side reconstruction quality comes from, with the omitted coordinates and storage cost held fixed so that only the fill source changes. Right: answer distortion and storage over four rollout-step ranges for teacher-forced oracle and closed-loop reconstruction.
Gate
∣ΔNLL∣
Rate
Rate ↑
All-send
0.000
3737.5
+58.8%
MemoWM
1.536
2354.2
–
Static- e
1.536
2754.3
+17.0%
Uniform- V
1.536
2954.8
+25.5%
Random
1.536
3134.2
+33.1%
Uncertainty
1.536
3342.1
+42.0%
Table 3: Task-aware gating at matched distortion. Rate: bits/state; Rate ↑ : increase over MemoWM.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Subset
Coords.
Low- Vj
Unseen query
Random
Bottom 5%
51
0.212
0.323
37.4
Bottom 10%
102
0.221
0.375
38.8
Bottom 15%
154
0.231
0.423
38.3
Bottom 20%
205
0.232
0.565
39.8
Appendix
Table 4: Held-out checks of the two risk components. (a) Answer-NLL damage conditioned on prediction errors for the original queries and independently constructed unseen-query variants ( 10−3 bits/answer). (b) Frozen entropy lookup versus observed prediction error (%).