cs.CLSep 29, 2026

Learning What to Remember: Long-horizon Counterfactual Memory Optimization

Authors: Jiaming Tang, Mingyan Liu, Armin Sarabi

Organizations: Department of Electrical Engineering and Computer Science, University of Michigan

Abstract

Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

    May 1, 2026Derong Xu, Shuochen Liu, Pengfei Luo +8Personalization

  2. HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents

    Jun 15, 2026Jiangze Yan, Yi Shen, Wenjing Zhang +5Long-Horizon AgentsLong-Term Agent Memory