Use and Disuse: Intent-Structured Experience Consolidation for Memory and Learning in LLM Agents
Authors: Xiangyi Zeng, Baihang Liu, Xutong Wang, Ze Jin, Yunpeng Li, Qixu Liu
Organizations: Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
The evolution of Large Language Model agents from single-task execution to long-term autonomous operation highlights the critical challenge of transforming continuous experiences into reusable knowledge. To address this, we propose Hippocam, a hierarchical memory and continual learning architecture. Hippocam draws inspiration from two characteristics of human memory: cognitive processes selectively maintain information relevant to current goals, while long-term memories form gradually through repeated consolidation. Accordingly, Hippocam structures an agent's ongoing work as nested intents. The active context remains centered on the current intent, while completed intents are consolidated into the task-relevant outcomes and state needed for subsequent work, rather than carrying forward their full working details. Concurrently, a recursive prefix consolidation mechanism repeatedly consolidates earlier history, causing long-unused experiences to become increasingly abstract. Original interactions are preserved, allowing the agent to progressively recover finer-grained details through the hierarchy and stop once sufficient information is available. Crucially, when past experiences are recalled and reintegrated into active work, they undergo subsequent consolidation alongside new experiences, thereby being reinforced, supplemented, and updated. Through this memory dynamic of use and disuse, Hippocam connects working context, long-term memory, knowledge accumulation, and skill learning within a single continuously evolving experiential process. This enables agents to learn and evolve capabilities through their own experiences without parameter updates.
Figures & tables
Figure 1: Hippocam unifies memory, knowledge use, and learning. Performance (%) with DeepSeek-V4.1-Flash as the task backbone. Evaluation details appear in § 4 .
Figure 2: Overview of Hippocam.
Figure 3: How Hippocam learns. Knowledge: reused ; new ; fading .
Method
ALF.
Sci.
Stream.
Plain
78.4
66.9
69.0
HiAgent
85.1
56.7
69.0
mem0
87.3
58.6
69.3
A-Mem
95.5
87.2
79.0
Letta
85.1
71.2
81.3
ReasoningBank
85.1
66.9
68.3
Table 1: Task performance (%) on ALF World , Sci enceWorld , and Stream Bench .
Figure 4: Performance as experience accumulates. Cumulative accuracy on StreamBench. Each stream’s top three methods at task 150 continue to task 300. Shading marks this extension.
Table 2: Cross-domain knowledge generalization and integration. ScienceWorld mean progress and FEVER accuracy (%). Δstream reports After minus Fresh in percentage points (pp).
Figure 5: Maintaining usable experience. Left: GoodAI LTM scores under isolated and 32k-span interleaved conditions. Right: MemoryAgentBench multi-hop conflict-resolution accuracy across input lengths. Lines connect separate length conditions.
Method
Overall
Single-hop
Multi-hop
Temporal
Open-domain
Adversarial
HiAgent
39.8
37.1
24.8
16.2
45.8
70.2
mem0
58.0
67.2
37.6
8.1
77.1
85.2
A-Mem
67.5
74.3
34.4
67.6
61.5
76.7
Letta
51.7
48.5
22.0
35.2
45.8
89.5
ReasoningBank
18.3
2.9
0.0
2.8
7.3
72.4
ACE
22.5
0.0
0.0
0.0
0.0
100.0
Table 3: What the retained conversation supports. Binary judge accuracy (%) on LoCoMo. Evaluation scope is specified in § A.3 .
Setting
Metric
Hippocam
w/o
Intent conditioning
GoodAI isolated
Score ↑
85.1
77.8
interleaved
Score ↑
88.7
65.6
Prefix consolidation
ToolBench
Acc. ↑
78.7
65.7
ALFWorld
Success ↑
97.8
100.0
Table 4: Component ablations and consolidation comparison. Scores, success rates, and accuracy are percentages. Context size is measured in tokens.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Default
Role
K
2
Minimum working-tail interaction rounds
Bmin
4,096 tokens
Working-tail token floor
α
0.20
Working-tail token cap as a fraction of W
q
0.95
Action-round length quantile
Emin
2
New intent consolidations in a stable prefix
Tmin
16,384 tokens
Stable-prefix size for ordinary consolidation
Appendix
Table 5: Default consolidation configuration. Token quantities are implementation budgets; W is the configured context capacity.
First 150 tasks per stream
First 300 tasks per stream
Method
DDXPlus
ToolBench
Mean
DDXPlus
ToolBench
Mean
Plain
74.7
63.3
69.0
—
—
—
HiAgent
76.7
61.3
69.0
—
61.3
—
mem0
70.0
68.7
69.3
—
69.0
—
A-Mem
84.0
74.0
79.0
91.3
74.3
82.8
Letta
86.0
76.7
81.3
91.7
77.3
84.5
Appendix
Table 6: StreamBench accuracy at two evaluation horizons. Mean gives equal weight to the two streams and is reported only when both reach the stated horizon.
ScienceWorld
FEVER
Method
Progress
Cost
Accuracy
Cost
HiAgent
42.7
2.62
61.8
1.76
mem0
51.7
0.06
59.9
0.25
A-Mem
42.1
0.26
55.8
0.35
Letta
40.8
2.53
59.9
0.98
ReasoningBank
53.3
0.10
55.8
0.19
Appendix
Table 7: Frozen-source performance and recorded target-evaluation cost. Performance is reported in percentages; cost is the total USD expenditure for the reported target evaluation, including within-task memory operations and excluding source experience acquisition.
Method
6k
32k
64k
Overall
DetectiveQA
HiAgent
97.0
0.0
11.0
36.0
—
Hippocam
97.0
77.0
41.0
71.7
85.9
Letta
77.0
37.0
19.0
44.3
82.4
A-Mem
45.0
17.0
11.0
24.3
78.9
mem0
19.0
25.0
17.0
20.3
—
ACE
0.0
0.0
0.0
0.0
—
Appendix
Table 8: MemoryAgentBench conflict resolution and long-range understanding (LRU). Accuracy (%); Overall averages the three multi-hop input lengths.
Configuration
Success (%)
Mean steps
Hippocam
100.0
8.12
Empty placeholder
94.0
9.92
ACE playbook
97.8
9.00
Flat summary
100.0
9.68
Appendix
Table 9: Prefix consolidation and alternative knowledge configurations. ALFWorld success and mean actions. The ACE variant replaces Prefix Consolidation with its playbook-generation process.
Figure 6: Retained detail for a later action. Local archive subtrees redrawn in the style of the Hippocam viewer, with verbatim memory and response excerpts. The left intent trace shows the completed shopping-list intent and the new customer intent, with labels shortened for display; node IDs identify saved records and ellipses mark omissions. The selected memories retain different levels of menu detail.
Intervening use
Unchanged
Supplemented
Updated
Rewritten
Dropped
No detected use
95.8
0.9
0.1
3.1
0.1
No API, unclassified
94.1
2.1
0.3
3.4
0.1
Used, call matches reference
57.0
28.4
0.3
14.3
0.0
Used, reference call differs
68.3
11.7
3.4
16.6
0.0
Appendix
Table 10: Knowledge-item transitions in ToolBench grouped by API-linked use between consecutive consolidations. Outcome columns report percentages within each group.
Environment
Textual reuse
Unchanged (%)
ALFWorld
Detected
41.6
Not detected
63.2
ScienceWorld
Detected
57.7
Not detected
59.1
Appendix
Table 11: Knowledge-item transitions grouped by four-word textual reuse in main-agent messages. Unchanged proportions are percentages.
Figure 7: Experience changes through use and disuse. Knowledge gains supporting evidence as episode-specific narrative detail recedes. Abridged and paraphrased from ALFWorld traces.
Figure 8: Progressive recall in two MemoryAgentBench conflict-resolution cases. Blue boxes are opened memory; dashed gray boxes are references or children left unopened. The second case preserves the observed reconsolidation and completion check between recall and answer.
Figure 9: Performance, reuse, and the cost of continued work. Top: accuracy versus average cost per task on each stream’s first 150 tasks; upper-left is preferable. The blue arrow connects Hippocam’s first 150 tasks (open marker) to its next 150 (filled marker); each endpoint reports its own interval. Bottom: the first-150 cost decomposed into cached input, uncached input, output, and embeddings. All components include memory maintenance.
Stream
Tasks
Acc. (%)
Input (k)
Uncached (k)
USD/task
USD/correct
DDXPlus
First 150
87.3
207.8
8.0
0.0095
0.0109
Next 150
98.7
252.3
8.3
0.0085
0.0087
ToolBench
First 150
76.0
232.5
30.4
0.0132
0.0173
Next 150
81.3
259.5
26.8
0.0118
0.0145
Appendix
Table 12: Later work improves without a higher mean cost per task. Hippocam’s consecutive StreamBench intervals. Input counts sum all agent calls per task, in thousands of tokens; they are not single-request context lengths. Cost per correct answer divides total cost, including incorrect answers, by the number correct.
Stream
Method
Acc. (%)
Input (k)
Uncached (k)
USD/task
USD/correct
DDXPlus
Hippocam
93.0
230.1
8.1
0.0090
0.0097
A-Mem
91.3
20.1
19.1
0.0083
0.0091
Letta
91.7
77.1
—
—
—
ToolBench
Hippocam
78.7
246.0
28.6
0.0125
0.0159
A-Mem
74.3
61.9
51.6
0.0193
0.0260
mem0
69.0
48.3
38.0
0.0470
0.0681
Appendix
Table 13: Cumulative cost–performance after the next 150 tasks. Scores and costs cover both intervals and include memory maintenance. Letta’s total input is recorded, but its uncached input and exact repriced cost cannot be recovered.
Large language model agents are expected to continuously adapt to new tasks and environments over their lifetime by reusing past experience. However, existing memory-based agents struggle to transfer reusable experience across environments and suffer from catastrophic forgetting as experience accumulated. To address these challenges, we propose LifeMem, a lifelong learning framework that enables agents to transfer knowledge across multiple environments. During learning, LifeMem clusters accumulated interaction trajectories based on underlying workflows to extract reusable skills. When solving a new task at inference time, the agent recalls relevant skills and trajectories to guide actions. To validate our method, we conduct experiments across 10 environments and over 13k tasks with 2k newly annotated interaction trajectories. Results show that LifeMem enables effective experience reuse in lifelong learning, achieving both reduced forgetting on learned tasks and superior cross-task transfer. Further analysis reveals that task streaming impacts learning, while consolidating structurally similar trajectories within memory boosts performance.
Yuli Qiu, Yutong Li, Wei Su +5
School of Computer Science and Technology, Beijing Institute of Technology · School of Computer Science and Engineering, Beihang University · Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology +1
Large Language Model (LLM)-based agents increasingly rely on memory to learn from experiences over continual interactions. However, storing experiences as independent, flat units leads to substantial redundancy and retrieval conflicts, as similar episodes repeat overlapping content and subtle scene variations cause retrieved memories to offer contradictory guidance. To address this, we introduce residual experience, positing that newly acquired experience is often an incremental variation of existing knowledge. We propose DeltaMem, a framework that organizes experience memory into two independent residual trees, one storing goal-conditioned task experience as reusable skills and another for scene-level environment knowledge. Each tree uses a root node for generalized base experiences and incremental delta nodes for subsequent variations, allowing related experiences to share a common foundation without duplication. For retrieval, a failure-penalized similarity scan locates the best match, reconstructing the full experience via root-to-match chain composition. An autonomous consolidation mechanism distills high-frequency paths into new root nodes, enabling the trees to self-organize from general heuristics to specialized variants. Experiments across diverse interactive environments show that DeltaMem consistently outperforms existing baselines. To facilitate future research, we release the code at https://github.com/import-myself/DeltaMem.
Haoran Tan, Zeyu Zhang, Zhicheng Cao +2
Gaoling School of Artificial, Renmin University of China, Beijing, China · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE +1
Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusing knowledge. Prior memory approaches aim to resolve the situation, but mainly focus on storing factual information. Recent work on procedural memory improves task reuse, yet often reduces to replaying past successes without addressing failure cases or online scalability. We introduce a unified and automatic memory framework that integrates semantic, episodic, and procedural memory in a bi-level design combining short-term and long-term stores. A multi-agent architecture with actor, memory, and critic agents enables automatic memory generation, reward annotation, and adaptive retrieval. Long-term memory is managed through reward-based evaluation, merging, and pruning, ensuring scalability and continual improvement. Experiments across various environments show that our approach improves robustness and success on long multi-turn tasks compared to existing baselines. This work highlights the importance of comprehensive, adaptive memory for advancing LLM-based agents.