Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. \textbf{In-parameter memory} offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: \textbf{Parameter Placement}, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and \textbf{Parameter Acquisition Time}, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.
Figures & tables
Figure 1: A demonstration of the representative methods in the two-dimensional taxonomy.
Figure 2: A demonstration of the operator formulation of in-parameter memory.
Paradigm
Parametric
Independent Acquisition
Augmented at Deployment
Included
ICL-based RAG
✘
✘
✓
✘
Pluggable PEFT
✓
✓
✓
✓
Full Parameter SFT
✓
✓
✘
✘
Unmodified KV Cache
✓
✘
✓
✘
Table 1: Scope relative to neighboring paradigms. Parametric asks whether the memory takes continuous parameter form. Independent acquisition asks whether it is produced by an independent acquisition operator A rather than as a byproduct of running the model. Augmented at deployment asks whether it enters the forward pass at deployment as a distinct object. A paradigm is included only when all three hold.
Dimension
Class
Operational Criterion
Typical Memory Object
Parameter Acquisition Time
Online
ϕ is generated or updated while the model is serving.
Fast Weights; Online Neural Memory; Adapters; Soft Tokens.
Offline
ϕ is formed before serving and held fixed, but still augmented at deployment.
Adapters; Memory Tables; Fast Weights; Online Neural Memory.
Hybrid
ϕ is composed at two or more of these sites.
Adapters, Online Neural Memory.
Table 2: The operational criteria and typical memory objects of different classes of methods in the two dimensions used to locate included methods.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Placement
Memory object
Acquisition operator A
Embedding — Online
MemGen ( Zhang et al., 2026b )
Embedding
Soft Tokens
RL-trained trigger + LoRA weaver
MINT ( Yi et al., 2025 )
Embedding
Soft Prompts
Test-time bank update
GradMem ( Kuratov et al., 2026 )
Embedding
Soft Tokens
Test-time gradient descent
REFRAG ( Lin et al., 2025b )
Embedding
Soft Tokens
Selective expansion policy
LatentMem ( Fu et al., 2026a )
Embedding
Soft Tokens
LMPO-trained composer
Appendix
Table 3: Memory object and acquisition operator of every method in Figure 3 . Memory objects use the seven types of Table 2 . All rows are core; permanently merged comparators are excluded.
Method
Task purpose
Material.
Persist.
Write
Read
Embedding — Online
MemGen ( Zhang et al., 2026b )
Agent self-evolution
Per-request
Ephem.
M
L
MINT ( Yi et al., 2025 )
Task adaptation
Per-request
Ephem.
M
L
GradMem ( Kuratov et al., 2026 )
Long-context modeling
Per-request
Ephem.
M
L
REFRAG ( Lin et al., 2025b )
Knowledge injection
Per-request
Ephem.
M
L
LatentMem ( Fu et al., 2026a )
Multi-agent Communication
Per-request
Ephem.
M
L
Appendix
Table 4: Deployment properties of the methods in Table 3 . Write/read cost are qualitative bands (L/M/H) inferred from the mechanism.
Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessary yet remains challenging. Although inference-time methods like In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG) are popular, they face constraints in context budgets, costs, and retrieval fragmentation. Departing from these context-dependent paradigms, this work investigates a parametric approach using Low-Rank Adaptation (LoRA) as a modular knowledge memory. Although few recent works examine this concept, the fundamental mechanics governing its capacity and composability remain largely unexplored. We bridge this gap through the first systematic empirical study mapping the design space of LoRA-based memory, ranging from characterizing storage capacity and optimizing internalization to scaling multi-module systems and evaluating long-context reasoning. Rather than proposing a single architecture, we provide practical guidance on the operational boundaries of LoRA memory. Overall, our findings position LoRA as the complementary axis of memory alongside RAG and ICL, offering distinct advantages.
Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distributed across past observations and actions. However, useful information encoded in previously computed representations is often underutilized during subsequent generation. We propose \textbf{TransMem}, a lightweight inference-time parametric memory module that transforms sparse historical hidden states from a frozen LLM backbone into reusable memory representations. TransMem uses a lightweight gating network to dynamically apply the latent intervention to the current hidden states, without repeatedly encoding the preceding context. To learn transferable memory utilization rather than task-specific knowledge, we introduce evidence-conditioned self-distillation. A memory-augmented student processes the full context and matches the predictive distribution of an evidence-only teacher that shares the same frozen backbone. Experiments on LoCoMo, HotpotQA, and MemoryAgentBench demonstrate consistent improvements across different model architectures and scales. TransMem yields gains of 11.58--29.25 F1 on LoCoMo and 10.20--13.03 F1 on HotpotQA, while improving the average MemoryAgentBench accuracy from 29.54% to 40.00%. These results establish sparse historical hidden states as an effective and efficient memory substrate for long-context LLM agents. Our code is available at https://github.com/Haodong-Lei-Ray/TransMem.
Haodong Lei, Junming Liu, Yirong Chen +4
Southeast University, Nanjing, Jiangsu, China · Shanghai Artificial Intelligence Laboratory, Shanghai, China
Long-running LLM agents require memory that persists and evolves across sessions. Text-based memory retrieves and reconstructs past interactions at every query, making long-horizon performance increasingly dependent on retrieval quality and contextual reasoning as histories grow. Parametric memory encodes experience directly into model computation, but existing approaches provide limited support for cross-session memory evolution. Their coupling to a specific backbone further restricts memory reuse after model replacement. We introduce RPMem, a two-stage architecture that compiles each session into a model-independent latent memory through forward computation and selectively integrates it with retained memory via a task-trained recurrent gate. The consolidated memory is then mapped to backbone-specific low-rank adaptation (LoRA) parameters, allowing the encoding capability to transfer when the backbone is replaced. Evaluation across three long-term memory benchmarks and five diverse backbones demonstrates broad generalization with near-constant update cost and memory footprint. With Qwen3-8B on PERMA, RPMem reaches 85.52%, outperforming the strongest parametric and text-based baselines by 5.32 and 12.98 percentage points, respectively. Ablations validate the complementary roles of session compilation and cross-session consolidation, while dynamics analyses reveal that the gate acquires task-specific memory integration strategies. These results establish RPMem as a lifecycle-independent parametric memory framework that maintains evolving cross-session memory that remains reusable across backbone replacements. Our implementation is available at https://github.com/Quark-Medical/rpmem/tree/main.