Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical applications, where agents must incorporate domain-specific knowledge and adapt from experience. Context engineering offers a practical alternative to weight updates by improving model behavior through instructions, strategies, and evidence supplied at inference time. However, adapting context online typically requires a costly trial-and-error process, while queries are often processed independently, preventing useful experience from carrying forward. Memory systems address this limitation by retaining information across interactions, but approaches that continually append information to a shared context face increasing token costs, context-window limits, and performance degradation as the context expands. We introduce a unified formulation of context optimization and show that an agent memory system update can be interpreted as an optimization update procedure over the model's context. This perspective attempts to provide a principled framework for studying memory design and its efficiency. We then propose GraphMemory, a lightweight graph-based memory that accumulates, refines, organizes, and connects reusable strategies. For each query, GraphMemory retrieves only the relevant subgraph, enabling online context adaptation without exposing the model to the entire memory. Under bounded retrieval, the amount of retrieved memory remains constant as the number of processed examples grows. Experiments show that GraphMemory achieves competitive downstream performance while using approximately 81-85% fewer memory-construction tokens than our baselines.
Figures & tables
Figure 1: Prompt-token usage at successive training checkpoints for GraphMemory and ACE on Formula with Qwen3.8-27B.
Figure 2: Overview of GraphMemory. An LLM retriever selects entry nodes from a compact index, after which a bounded breadth-first traversal loads the relevant subgraph. The ACE Generator and Reflector operate on this retrieved context, while the Curator writes new nodes and edges back to the graph.
Dataset
Backbone
System
Accuracy
Token Usage
Estimated Cost *
Initial
Best val. (step)
Final test
Training
Test
Training (USD)
Formula
Claude Haiku 4.5
GraphMemory
0.70
0.87 (500)
0.81
∼ 20M
∼ 1.4M
$24.00
ACE (80k cap)
0.70
0.88 (300)
0.85
∼ 131M
∼ 18M
$157.20
Qwen3.8-27B
GraphMemory
0.69
0.88 (200)
0.84
∼ 8M
∼ 0.6M
$5.00
ACE
0.67
0.88 (200)
0.83
∼ 42M
∼ 4M
$26.25
FiNER
Claude Haiku 4.5
GraphMemory
0.75
0.80 (700)
0.75
∼ 57M
∼ 9M
$68.40
Table 1: Offline adaptation performance, token usage, and estimated training cost on Formula and FiNER . Training-token counts and costs are approximate. Costs assume a 95%−5% input–output token split and standard real-time API prices. Initial accuracy denotes cold-start test accuracy. † ACE/Qwen on FiNER is not reported; see note below.
Figure 3: Validation performance of GraphMemory and ACE on Claude across training steps on FiNER (left) and Formula (right).
Dataset
Model
System
Training
Final test
Total
Per step
Total
Per sample
Formula
claude-haiku-4-5
GraphMemory
6.4 h
46 s
15.7 min
4.7 s
ACE (80k cap)
8.1 h
58 s
8.6 min
2.6 s
FiNER
claude-haiku-4-5
GraphMemory
19.0 h
68 s
47.1 min
6.4 s
ACE (80k cap)
28.9 h
104 s
26.5 min
3.6 s
Table 2: Active wall-clock time for offline training and held-out test on Formula and FiNER . All figures are active wall clock. Bold indicates the faster result.
Long-context language modeling requires not only extending context windows but maintaining coherent understanding of entity states and relationships across thousands of tokens -- a challenge that semantic similarity alone cannot address. KGERMAR addresses this by constructing dynamic, context-specific knowledge graphs from input text during inference, enabling domain-adaptive retrieval that leverages both semantic similarity and explicit entity relationships. The framework performs real-time entity and relation extraction to build contextual knowledge graphs, then integrates graph-structural embeddings with textual semantics through a multi-component memory architecture. Three memory banks -- contextual, semantic, and structural -- are maintained with retrieval signals fused via learned weights to capture both surface-level semantics and deeper relational patterns. Evaluated on SlimPajama (84.7K training examples), WikiText-103 (4,358 examples), PG-19 (100 examples), and Proof-pile (46.3K examples), KGERMAR achieves up to 8.5% lower perplexity and 2--2.5x better memory efficiency than memory-augmented baselines across context lengths from 1K to 32K tokens, with superior in-context learning performance across five NLU tasks. The dynamic knowledge graph construction approach advances memory-augmented language modeling by enabling domain-specific knowledge representation that adapts to input contexts rather than relying on fixed knowledge bases.
Ghadir Alselwi, Basem Suleiman, Hao Xue +4
University of New South Wales, Sydney, NSW, Australia · 2Hong Kong University of Science and Technology (Guangzhou) · University of Southampton, Southampton, UK +2
Memory is an indispensable capability for long-horizon LLM agents, enabling them to preserve and utilize information accumulated across extended interactions. Existing memory-agent approaches are typically trained end-to-end with reinforcement learning on downstream tasks. However, collecting high-quality annotated problems for memory-intensive scenarios is costly, and the resulting training data often lack sufficient diversity to cover general memory behaviors. In this work, we propose MemTrain, a self-supervised training framework for generally enhancing the context-memory capability of LLM agents for more effective downstream post-training. MemTrain introduces two coupled proxy tasks over unlabeled Wikipedia corpora: (1) an end-to-end masked reconstruction objective, which requires the model to recover masked entities after multiple rounds of memory updates, thereby encouraging memory maintenance from the final outcome perspective; and (2) an intermediate memory recall objective, which requires the model to reconstruct masked historical information using intermediate memory states, encouraging faithful compression and memory completeness throughout the interaction process. The two objectives are jointly optimized using GRPO. Extensive experiments on long-text QA and search-based QA benchmarks demonstrate that MemTrain consistently improves downstream memory-intensive reasoning performance across different models, achieving gains of up to 17.67 points over direct task-specific post-training.
Ziheng Li, Xingrun Xing, Haoqing Wang +2
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Samsung Research, Beijing, China
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu +11
University of California Los Angeles · University of Washington · Northwestern University