Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. \textbf{In-parameter memory} offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: \textbf{Parameter Placement}, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and \textbf{Parameter Acquisition Time}, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.
Figures & tables
Figure 1: A demonstration of the representative methods in the two-dimensional taxonomy.
Figure 2: A demonstration of the operator formulation of in-parameter memory.
Paradigm
Parametric
Independent Acquisition
Augmented at Deployment
Included
ICL-based RAG
✘
✘
✓
✘
Pluggable PEFT
✓
✓
✓
✓
Full Parameter SFT
✓
✓
✘
✘
Unmodified KV Cache
✓
✘
✓
✘
Table 1: Scope relative to neighboring paradigms. Parametric asks whether the memory takes continuous parameter form. Independent acquisition asks whether it is produced by an independent acquisition operator A rather than as a byproduct of running the model. Augmented at deployment asks whether it enters the forward pass at deployment as a distinct object. A paradigm is included only when all three hold.
Dimension
Class
Operational Criterion
Typical Memory Object
Parameter Acquisition Time
Online
ϕ is generated or updated while the model is serving.
Fast Weights; Online Neural Memory; Adapters; Soft Tokens.
Offline
ϕ is formed before serving and held fixed, but still augmented at deployment.
Adapters; Memory Tables; Fast Weights; Online Neural Memory.
Hybrid
ϕ is composed at two or more of these sites.
Adapters, Online Neural Memory.
Table 2: The operational criteria and typical memory objects of different classes of methods in the two dimensions used to locate included methods.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Placement
Memory object
Acquisition operator A
Embedding — Online
MemGen ( Zhang et al., 2026b )
Embedding
Soft Tokens
RL-trained trigger + LoRA weaver
MINT ( Yi et al., 2025 )
Embedding
Soft Prompts
Test-time bank update
GradMem ( Kuratov et al., 2026 )
Embedding
Soft Tokens
Test-time gradient descent
REFRAG ( Lin et al., 2025b )
Embedding
Soft Tokens
Selective expansion policy
LatentMem ( Fu et al., 2026a )
Embedding
Soft Tokens
LMPO-trained composer
Appendix
Table 3: Memory object and acquisition operator of every method in Figure 3 . Memory objects use the seven types of Table 2 . All rows are core; permanently merged comparators are excluded.
Method
Task purpose
Material.
Persist.
Write
Read
Embedding — Online
MemGen ( Zhang et al., 2026b )
Agent self-evolution
Per-request
Ephem.
M
L
MINT ( Yi et al., 2025 )
Task adaptation
Per-request
Ephem.
M
L
GradMem ( Kuratov et al., 2026 )
Long-context modeling
Per-request
Ephem.
M
L
REFRAG ( Lin et al., 2025b )
Knowledge injection
Per-request
Ephem.
M
L
LatentMem ( Fu et al., 2026a )
Multi-agent Communication
Per-request
Ephem.
M
L
Appendix
Table 4: Deployment properties of the methods in Table 3 . Write/read cost are qualitative bands (L/M/H) inferred from the mechanism.