Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
Figures & tables
Figure 1: Overview of FactorEngram. Left: memory modules are residually inserted before the attention module at selected Transformer layers. Right: n-gram lookups retrieve coefficients that are concatenated and modulated by context-dependent basis-level gates. The shared dictionary participates in both gating and memory reconstruction. The reconstructed memory is projected, refined by a short causal convolution, and added to the backbone hidden state.
Model
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
340M backbone
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
Engram
22.29
23.94
49.06
69.8
37.6
15.0
FactorEngram
21.29
21.69
49.82
70.6
79.1
58.3
1B backbone
Table 1: Main results with an 8K context length. The 340M and 1B backbones are trained on 30B and 120B tokens, respectively. Accuracy metrics are reported as percentages. Bold indicates the best value within each model scale.
Configuration
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
Engram
22.29
23.94
49.06
69.8
37.6
15.0
Engram (+ unigram)
22.32
24.42
48.94
74.0
42.2
29.2
FactorEngram (scalar gate)
22.85
25.66
48.15
67.8
45.2
9.4
FactorEngram (- unigram)
22.74
23.76
49.20
60.2
44.2
23.0
Table 2: Architectural component ablations with the 340M backbone. All models use an 8K context length. Accuracy metrics are reported as percentages. Bold indicates the best value in each column.
Configuration
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
λ=0
21.47
22.58
49.49
58.8
70.6
31.8
λ=10−4
21.18
22.43
49.43
67.4
74.0
43.2
λ=10−3
21.29
21.69
49.82
70.6
79.1
58.3
λ=10−2
23.89
25.79
47.77
50.0
50.0
17.8
Table 3: Effect of the sparsity coefficient λ with the 340M backbone and an 8K context length. Accuracy metrics are reported as percentages. Bold indicates the best value in each column.
Figure 2: Insertion-depth sweeps. The top row varies the single insertion layer and the bottom row fixes one insertion at optimal single layer 12 and varies the second. We report average downstream accuracy (%), perplexity, and NIAH accuracy (%). Colors distinguish metrics. Solid and dashed lines represent FactorEngram and Transformer, respectively. Dotted vertical lines mark the selected optimal depth. Lower perplexity and higher accuracy are better.
Insertion location
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
Before attention
21.29
21.69
49.82
70.6
79.1
58.3
Before FFN
21.28
22.04
48.40
65.8
63.8
18.8
Inside FFN
23.34
26.91
47.90
47.4
47.4
12.2
Table 4: Comparison of memory insertion locations with the 340M backbone and an 8K context length. Accuracy metrics are reported as percentages. Bold indicates the best value in each column.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Memory Representation
Memory Content
Factorization
Fine-Grained Gate
Sparse Coding
Unigram ( N=1 )
N-gram ( N≥2 )
Engram
✓
TN-gram
✓
✓
Lngram
✓ ∗
Engram-Nine
✓
Tokenizer-Agnostic Engram
✓
✓
Appendix
Table 5: Comparison of Engram and its variants. Check marks indicate features reported in each main method. Factorization refers to parameterizing memory entries through shared factors. Fine-grained gating refers to assigning component-wise contextual weights to a memory representation. Sparse coding refers to explicitly encouraging sparsity within representations through constraints or regularization. Memory content excludes backbone input token embeddings. ∗ Lngram retrieves N-grams of learned discrete latent symbols rather than input tokens.
Configuration
340M Backbone
1B Backbone
Total Params
1.4B
3.9B
Backbone Params
340M
1B
Total Tokens
30B
120B
Layers
24
Sequence Length
8192
Vocab Size
32K
Appendix
Table 6: Default FactorEngram configurations and training settings at the two backbone scales.
340M backbone
Category
Metric
Transformer
Engram
FactorEngram
Language Modeling
WikiText PPL ↓
23.19
22.29
21.29
LAMBADA PPL ↓
24.90
23.94
21.69
Downstream
LAMBADA ↑
37.32
38.15
39.30
PIQA ↑
67.22
68.23
68.06
HellaSwag ↑
43.05
44.00
46.14
Appendix
Table 7: Detailed main results with an 8K context length. We additionally report every downstream task accuracy (%).