Organizations: South China University of Technology · University of the Chinese Academy of Sciences · Beijing University of Posts and Telecommunications · University of Science and Technology of China · Tsinghua University · Peking University · Fudan University · Shanghai Jiao Tong University
Long-term agents face growing storage demands as they accumulate experience. World models capture reusable regularities that can reduce the information stored for each experience. We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content. Its task-aware allocation rule balances the expected impact of reconstruction errors against storage cost, retaining information with downstream value beyond the predictive prior. Across five long-term agent-memory benchmarks, MemoWM achieves 42.42% average answer accuracy, exceeding the strongest baseline by 2.62 percentage points, while reducing average experience-specific storage by 53.9% relative to MIRIX, the most storage-efficient baseline. Further analysis shows that stronger world models reduce per-experience storage at comparable task quality. Accounting for model parameters reveals a trade-off between shared model capacity and recurring storage costs, with the capacity that minimizes total storage increasing as more interactions are retained. Our code is available at https://github.com/Feld-maxiu/MemoWM.
Figures & tables
Figure 1: Average accuracy versus experience-specific storage.
Figure 2: Overview of the MemoWM framework. Coordinates are retained when their estimated task impact exceeds the weighted storage cost; others are filled by a causal world model. (A) Task-relevant retention, (B) shared-history reconstruction, and (C) offline fitting and calibration. Reconstructed states are retrieved through the Reader Bridge to answer future queries.
Method
WorldMemArena
AMA-Bench
LME-V2
WebQuest
MemEye
Avg.
RAG
52.97 ± 0.12
44.87 ± 0.70
40.25 ± 0.69
22.35 ± 0.50
35.42 ± 0.93
39.17
UniversalRAG
40.38 ± 0.19
40.35 ± 0.88
34.56 ± 1.18
21.68 ± 0.75
37.95 ± 1.24
34.98
MEM1
45.66 ± 0.31
14.96 ± 1.64
29.41 ± 1.83
20.25 ± 1.01
32.05 ± 1.93
28.47
MemAgent
43.03 ± 0.37
29.18 ± 1.60
29.33 ± 1.91
16.90 ± 0.78
30.88 ± 2.34
29.86
A-Mem
55.37 ± 0.28
34.85 ± 1.31
41.28 ± 1.48
26.70 ± 0.85
36.33 ± 1.28
38.91
MemGPT
58.43 ± 0.20
35.44 ± 0.81
44.97 ± 1.29
24.55 ± 0.87
35.61 ± 1.16
39.80
Table 1: Main results across five long-horizon memory benchmarks. (a) The upper table reports answer accuracy (%); (b) the middle table reports the experience-specific storage rate (kbit/observation); (c) the break-even point is the number of retained observations at which MemoWM’s total storage is no greater than the listed comparator’s. LME-V2 denotes LongMemEval-V2, Omni-SM denotes Omni-SimpleMem, and Avg. is the unweighted mean across benchmarks.
Figure 3: (a) Upper left: all-coordinate coding rate and (b) upper right: gated total rate with amortized FP16 weights, both versus model size at three training-data fractions. (c) Lower left: all-coordinate rate under four predictive contexts; blue curves show increases over full context (upper axis). (d) Lower right: storage-optimal gated model size versus retained transitions.
Figure 4: Left: where decoder-side reconstruction quality comes from, with the omitted coordinates and storage cost held fixed so that only the fill source changes. Right: answer distortion and storage over four rollout-step ranges for teacher-forced oracle and closed-loop reconstruction.
Gate
∣ΔNLL∣
Rate
Rate ↑
All-send
0.000
3737.5
+58.8%
MemoWM
1.536
2354.2
–
Static- e
1.536
2754.3
+17.0%
Uniform- V
1.536
2954.8
+25.5%
Random
1.536
3134.2
+33.1%
Uncertainty
1.536
3342.1
+42.0%
Table 3: Task-aware gating at matched distortion. Rate: bits/state; Rate ↑ : increase over MemoWM.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Subset
Coords.
Low- Vj
Unseen query
Random
Bottom 5%
51
0.212
0.323
37.4
Bottom 10%
102
0.221
0.375
38.8
Bottom 15%
154
0.231
0.423
38.3
Bottom 20%
205
0.232
0.565
39.8
Appendix
Table 4: Held-out checks of the two risk components. (a) Answer-NLL damage conditioned on prediction errors for the original queries and independently constructed unseen-query variants ( 10−3 bits/answer). (b) Frozen entropy lookup versus observed prediction error (%).
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can still omit task-critical facts, corrupt product attributes, or apply incorrect transition rules. To address such systematic prediction errors, we introduce MemWM, a memory-augmented text-based world model. MemWM uses world memory, a curated memory bank of transition rules, state caches, and hard-to-predict facts, to condition next-state imagination. We evaluate factual state preservation with Structured State Fidelity (SSF), which scores predicted states through benchmark-specific facts and fields. Compared with SFT, memory-augmented training improves SSF by up to 206.3%. In the full planning setting, we keep the policy model frozen and provide policy-side world skill: retrieved task-level skills and step-wise corrective guidance for action selection. Across ALFWorld, WebShop, and ScienceWorld, memory-augmented agents improve downstream success over an SFT-trained world-model agent, with up to a 65.4% relative gain. Sensitivity analyses further show that retrieved memory improves task success and efficiency under different memory and action-budget settings.
Yujun Wang, Tao Zhang, Jinhe Bi +9
LMU Munich · Munich Center for Machine Learning (MCML) · Huawei Heisenberg Research Center +4
We investigate when belief-based memory actually improves large language model (LLM) agents. Our vehicle is Nous, a long-term memory architecture that represents each entity-attribute pair as a categorical probability distribution updated through closed-form Bayesian inference, with information-theoretic surprise driving belief revision and entropy-based forgetting. A controlled ablation on the LoCoMo benchmark shows that Bayesian belief updating alone provides little benefit over naive last-write-wins because existing conversational memory benchmarks rarely contain contradictory or differently reliable evidence. We then introduce reliability-conditioned updating, estimating per-observation reliability from epistemic language, and show on a controlled contradiction benchmark that belief updating substantially outperforms last-write-wins and raw-memory retrieval when observations differ in trustworthiness. Because content-derived reliability is itself vulnerable to manipulation, we further propose provenance-capped belief updating, where trust is bounded by source provenance rather than textual confidence. Under controlled memory-poisoning experiments, this approach resists volumetric poisoning attacks while revealing the utility costs and implementation requirements of provenance-aware memory. Finally, we quantify a 27.5-point discrepancy between strict token-F1 and LLM-as-judge evaluation on identical outputs, highlighting important reproducibility concerns for long-term memory benchmarks. Our results suggest that probabilistic belief-based memory is most beneficial in environments requiring reasoning over conflicting and differently trustworthy evidence, rather than conventional conversational recall alone.
Large language model (LLM) agents are increasingly expected to operate over long-term interactions, where information from past dialogues must be preserved and recalled to support future tasks. However, as interactions accumulate, the memory store grows without bound and fills with redundant entries that inflate storage cost and degrade retrieval by crowding out the most useful evidence. Furthermore, this is especially limiting on resource-constrained platforms with hard memory budgets, motivating us to formulate storage-budgeted memory management, the task of keeping an already constructed memory store within a fixed budget while preserving information useful for future interactions. To this end, we then propose MemRefine, an LLM-guided framework that, since surface similarity poorly reflects factual value, uses similarity only to propose candidate pairs and defers delete, merge, and preserve decisions to an LLM judge based on factual content, iterating until the budget is met. Across multiple memory frameworks and long-term conversation benchmarks, MemRefine consistently meets target budgets while preserving downstream performance and outperforming rule-based baselines under tight budgets.