Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
Figures & tables
Figure 1: Overview of FactorEngram. Left: memory modules are residually inserted before the attention module at selected Transformer layers. Right: n-gram lookups retrieve coefficients that are concatenated and modulated by context-dependent basis-level gates. The shared dictionary participates in both gating and memory reconstruction. The reconstructed memory is projected, refined by a short causal convolution, and added to the backbone hidden state.
Model
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
340M backbone
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
Engram
22.29
23.94
49.06
69.8
37.6
15.0
FactorEngram
21.29
21.69
49.82
70.6
79.1
58.3
1B backbone
Table 1: Main results with an 8K context length. The 340M and 1B backbones are trained on 30B and 120B tokens, respectively. Accuracy metrics are reported as percentages. Bold indicates the best value within each model scale.
Configuration
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
Engram
22.29
23.94
49.06
69.8
37.6
15.0
Engram (+ unigram)
22.32
24.42
48.94
74.0
42.2
29.2
FactorEngram (scalar gate)
22.85
25.66
48.15
67.8
45.2
9.4
FactorEngram (- unigram)
22.74
23.76
49.20
60.2
44.2
23.0
Table 2: Architectural component ablations with the 340M backbone. All models use an 8K context length. Accuracy metrics are reported as percentages. Bold indicates the best value in each column.
Configuration
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
λ=0
21.47
22.58
49.49
58.8
70.6
31.8
λ=10−4
21.18
22.43
49.43
67.4
74.0
43.2
λ=10−3
21.29
21.69
49.82
70.6
79.1
58.3
λ=10−2
23.89
25.79
47.77
50.0
50.0
17.8
Table 3: Effect of the sparsity coefficient λ with the 340M backbone and an 8K context length. Accuracy metrics are reported as percentages. Bold indicates the best value in each column.
Figure 2: Insertion-depth sweeps. The top row varies the single insertion layer and the bottom row fixes one insertion at optimal single layer 12 and varies the second. We report average downstream accuracy (%), perplexity, and NIAH accuracy (%). Colors distinguish metrics. Solid and dashed lines represent FactorEngram and Transformer, respectively. Dotted vertical lines mark the selected optimal depth. Lower perplexity and higher accuracy are better.
Insertion location
Language Modeling PPL
Downstream
Long-context Retrieval
WikiText ↓
LAMBADA ↓
Avg. Acc. ↑
NIAH-1 ↑
NIAH-2 ↑
NIAH-3 ↑
Transformer
23.19
24.90
48.10
44.2
47.7
22.1
Before attention
21.29
21.69
49.82
70.6
79.1
58.3
Before FFN
21.28
22.04
48.40
65.8
63.8
18.8
Inside FFN
23.34
26.91
47.90
47.4
47.4
12.2
Table 4: Comparison of memory insertion locations with the 340M backbone and an 8K context length. Accuracy metrics are reported as percentages. Bold indicates the best value in each column.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Memory Representation
Memory Content
Factorization
Fine-Grained Gate
Sparse Coding
Unigram ( N=1 )
N-gram ( N≥2 )
Engram
✓
TN-gram
✓
✓
Lngram
✓ ∗
Engram-Nine
✓
Tokenizer-Agnostic Engram
✓
✓
Appendix
Table 5: Comparison of Engram and its variants. Check marks indicate features reported in each main method. Factorization refers to parameterizing memory entries through shared factors. Fine-grained gating refers to assigning component-wise contextual weights to a memory representation. Sparse coding refers to explicitly encouraging sparsity within representations through constraints or regularization. Memory content excludes backbone input token embeddings. ∗ Lngram retrieves N-grams of learned discrete latent symbols rather than input tokens.
Configuration
340M Backbone
1B Backbone
Total Params
1.4B
3.9B
Backbone Params
340M
1B
Total Tokens
30B
120B
Layers
24
Sequence Length
8192
Vocab Size
32K
Appendix
Table 6: Default FactorEngram configurations and training settings at the two backbone scales.
340M backbone
Category
Metric
Transformer
Engram
FactorEngram
Language Modeling
WikiText PPL ↓
23.19
22.29
21.29
LAMBADA PPL ↓
24.90
23.94
21.69
Downstream
LAMBADA ↑
37.32
38.15
39.30
PIQA ↑
67.22
68.23
68.06
HellaSwag ↑
43.05
44.00
46.14
Appendix
Table 7: Detailed main results with an 8K context length. We additionally report every downstream task accuracy (%).
Modern language models represent text using discrete token-level embeddings, which forces recurring multi-token patterns to be learned implicitly across Transformer layers. Both Over-tokenized Transformers and Engram attempt to address this limitation by explicitly incorporating multi-token (n-gram) memories. However, they rely on separate hash tables for each n-gram order, which introduces hash collisions and prevents nested n-grams from sharing the underlying latent structures. To address these issues, we propose Tensorized Engram (TN-gram), a compact memory module that represents tensorized n-gram embeddings through shared factors in the Canonical Polyadic (CP) form. TN-gram learns shared token-position factors together with order-absorption vectors to encode the embeddings of different n-gram order. Comprehensive experiments demonstrate that TN-gram matches or even outperforms Engram-style n-gram modules while requiring much fewer parameters.
Wuyang Zhou, Yuxuan Gu, Giorgos Iacovides +3
Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom · AIP, RIKEN, Tokyo, Japan.
Sequence modeling requires both compositional reasoning and local static knowledge retrieval, yet standard Transformers handle both through dense computation. Engram partially decouples retrieval from the backbone, but its token-based keys remain tied to text tokenization and hash compression. We propose Lngram, a latent-space conditional memory module that learns discrete symbols directly from hidden states and performs N-gram lookup over these symbols. This design removes the dependence on tokenizer IDs and naturally extends to non-text modalities. In our evaluated settings, Lngram outperforms Transformer and Engram baselines, consistently reduces perplexity in long-context language modeling, and effectively injects domain knowledge when added post hoc to pretrained models. Joint training with the backbone further surpasses full fine-tuning, while experiments on vision-language and vision-language-action tasks show overall gains. Analyses with LogitLens and CKA suggest that Lngram enables prediction-relevant information to emerge earlier, increasing effective depth with limited inference and memory overhead. Code is available at https://github.com/zyaaa-ux/Lngram.
Yunao Zheng, Guoyang Xia, Xiaojie Wang +1
Beijing University of Posts and Telecommunications (BUPT), Beijing, China · Li Auto Inc., Beijing, China
Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled with the backbone width, limiting scalability due to high parameter and activation costs. We propose Lngram v2, which decouples the number of routes, memory dimension, and backbone width, and introduces a context-aware grouped-query attention readout to scale memory capacity independently. A zero-value Sink and counterfactual surrogate gradients further improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments across vision--language models (VLMs) of different scales show consistent improvements, including successful scaling to a 30B-parameter model. Compared with Lngram v1, Lngram v2 substantially reduces both total and activated memory parameters while maintaining or improving language modeling performance. Further analysis shows that its discrete IDs preserve substantial semantic structure of continuous hidden states, enabling semantic recovery from IDs alone and stable ID--semantic associations across datasets. These results establish Lngram v2 as an efficient and scalable latent conditional memory mechanism whose discrete addresses also provide a structured interface for analyzing internal model representations.
Yunao Zheng, Bin Wen, Xiaojie Wang +12
Beijing University of Posts and Telecommunications. · Kuaishou Technology.