Frozen Memory Is Not Enough: Rethinking External Memory as Extraction
Organizations: ELLIS Institute Finland · University of Turku · University of Technology Sydney
Abstract
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.
Figures & tables
| Variant | Budget | NQ | WebQA | TriviaQA | TruthQA | HotpotQA | Average |
| Base | – | 20.6 | 29.3 | 57.7 | 32.1 | 21.0 | 32.1 |
| Non-parametric Methods | |||||||
| RAG [ Lewis et al., 2020 ] | – | 22.6 1.9 | 24.9 -4.4 | 54.2 -3.4 | 35.5 3.4 | 29.8 8.8 | 33.4 (+3.9%) |
| kNN [ Khandelwal et al., 2020 ] | – | 21.1 0.4 | 30.5 1.2 | 57.8 0.1 | 32.3 0.2 | 21.2 0.2 | 32.6 (+1.4%) |
| Parametric Methods | |||||||
| CPT | – | 12.2 -8.5 | 34.1 4.8 | 61.2 3.6 | 29.2 -2.9 | 16.0 -4.9 | 30.5 (-5.0%) |
| Variant | Budget | NQ | WebQA | TriviaQA | TruthQA | HotpotQA | Average |
| Base | – | 20.6 | 29.3 | 57.7 | 32.1 | 21.0 | 32.1 |
| Transfer Verification | |||||||
| Mistral -R4 | 10/20 | 30.2 9.5 | 34.0 4.7 | 70.0 12.3 | 30.8 -1.3 | 27.7 6.8 | 38.5 (+20.0%) |
| LLaMA -R4 | 10/20 | 30.3 9.7 | 33.7 4.4 | 69.9 12.3 | 30.9 -1.2 | 27.6 6.7 | 38.5 (+20.0%) |
| Frozen Mistral -R4 | 10/0 | 30.5 9.8 | 30.3 1.0 | 70.8 13.2 | 32.2 0.2 | 27.7 6.8 | 38.3 (+19.2%) |
| Frozen LLaMA -R4 | 10/0 | 30.6 10.0 | 30.3 1.0 | 70.8 13.1 | 32.2 0.1 | 27.6 6.7 | 38.3 (+19.2%) |
| Task | Transferred | Random | Disabled | Ablated | (Trans. - Random) | (Trans. - Disabled) | (Trans. - Ablated) |
| NQ | 25.1 | 20.1 | 18.3 | 26.2 | +0.035 | +0.032 | -0.013 |
| WebQA | 32.3 | 27.4 | 31.1 | 33.6 | +0.147 | +0.197 | -0.001 |
| TriviaQA | 72.5 | 65.5 | 58.9 | 72.5 | +0.047 | +0.072 | -0.000 |
| TruthQA | 30.8 | 32.6 | 31.7 | 30.9 | -0.581 | -0.258 | -0.029 |
| HotpotQA | 27.1 | 22.9 | 18.6 | 27.2 | +0.022 | +0.019 | -0.009 |
| Condition | Trainable Params | 5M | 20M | 50M |
| Transferred | 1.05M | |||
| From scratch | 34.6M |
| Condition | Test PPL | 95% CI | NQ | WebQA | TriviaQA | TruthQA | HotpotQA | Avg |
| No memory baseline | ||||||||
| Transferred ( ) | ||||||||
| Interface Simplifications | ||||||||
| No gate | ||||||||
| Affine stitch | ||||||||
| Content and Training Controls | ||||||||
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
| Source model and target model, respectively. | |
| Hidden dimensions of source model and target model , respectively. | |
| Frozen Engram memory artifact learned with source model , comprising the tables . | |
| Memory table associated with N-gram order and hash head , where . | |
| Maximum N-gram order used for memory addressing, with . | |
| Number of independent hash heads for each N-gram order. |
| Property | KNN-LM | RAG | RETRO | Mem. Layers | LLM Mod. | Engram | Ours |
| retrieval | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ |
| Parametric (trained) | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Cross-model portable | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | |
| Surgical deletion | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ |
| No context overhead | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | |
| Reader-only integration | N/A | N/A | ✗ | ✗ | ✗ | ✗ | ✓ |
| Aspect | Original Engram | Transfer-oriented Engram Reader |
| Memory indexing | Layer-specific compressed-token -gram hashing: | Shared canonical hashing: |
| Memory representation | Layer-specific Engram embedding: | Shared memory vector: |
| Key / gate | ||
| Output update | ||
| Branch structure | Native multi-branch inside backbone | Explicit reader branches ( controls capacity) |
| Injection | Internal transformer block component | Post-layer hook injection (e.g., layers 2 and 10) |
| Family | Model | Experimental role | Params. | Checkpoint | Hugging Face identifier | |
| Pythia | Pythia-160M | Source: main matrix | 160M | 768 | Base | EleutherAI/pythia-160m |
| Pythia-410M | Target: main matrix | 410M | 1024 | Base | EleutherAI/pythia-410m | |
| Qwen3.5 | Qwen3.5-0.8B | Source: main matrix, scaling Target: downstream | 0.8B | 1024 | Base | Qwen/Qwen3.5-0.8B-Base |
| Qwen3.5-2B | Target: scaling, downstream | 2B | 2048 | Base | Qwen/Qwen3.5-2B-Base | |
| Qwen3.5-4B | Source: peer, downstream Target: main matrix, peer, scaling, downstream | 4B | 2560 | Base | Qwen/Qwen3.5-4B-Base | |
| Qwen3.5-9B | Source: main matrix, downstream Target: scaling, downstream | 9B | 4096 | Base | Qwen/Qwen3.5-9B-Base |
| Regime / phase | Corpus | Budget / cap | Seq. / batch | LR / warm-up | Trainable components and exceptions |
| Primary training settings | |||||
| RQ1 4.1 Phase 1 | WikiText-103 | 50M tokens | 512 / 16 | 1,000 | Source backbone, memory, and source reader are trainable. For Qwen3.5 sources, the memory learning rate is . |
| RQ1 4.1 Phase 2 | WikiText-103 | 20M tokens | 512 / 16 | 500 | Only the target reader is trainable; the target backbone and transferred memory remain frozen. |
| Figure 3 Phase 1 | FineWeb-Edu | 50M tokens | 512 / 2 | / 1,000 | Source backbone, memory, and source reader trained end-to-end; gradient checkpointing |
| Figure 3 Phase 2 | FineWeb-Edu | 20M tokens | 512 / 4 | / 500 | Target reader only; target backbone and transferred memory frozen; gradient checkpointing; early stopping patience 5 |
| Table 1 to Table 5 Phase 1 | Wikipedia-2021 | 10–30M tokens | 2048 / 1 | 1,000 | The source memory and reader are trainable; the source backbone is frozen. Early stopping uses patience 3. |
| Target | Baseline | Transferred | Random |
| Pythia-160M source | |||
| Pythia-410M | -1.6% | ||
| Qwen3.5-4B | -6.8% | ||
| TinyLlama-1.1B | -10.6% | ||
| Qwen3.5-0.8B source | |||
| Pythia-410M | -6.8% | ||
| Condition | Test PPL | 95% CI | Params / Cost |
| Baseline (no memory) | — | ||
| Transfer Conditions | |||
| Random memory | -0.2% | 1.05M | |
| Transferred memory | -1.6% | 1.05M | |
| Parameter-Matched Baselines | |||
| LoRA (rank 7) [ Hu et al., 2022 ] | +6.5% | 1.03M | |
| Condition | Test PPL | 95% CI | Params |
| Baseline (no memory) | — | ||
| Transfer Conditions | |||
| Random memory | -5.9% | 2.10M | |
| Transferred memory | -10.6% | 2.10M | |
| Direction | Baseline | Transferred | Random |
| Phi-4-mini Qwen3.5-4B | -10.1% | ||
| Qwen3.5-4B Phi-4-mini | -9.2% |
| Target | Baseline | Transferred | Random |
| Qwen3.5-2B | -14.1% | ||
| Qwen3.5-4B | -8.9% | ||
| Qwen3.5-9B | -8.3% |
| Target | Source | RTE | BoolQ | OpenBookQA | SciQ | TruthfulQA | RACE |
| Qwen-0.8B | FW-avg | +6.1 | -0.6 | +0.7 | +1.8 | -0.1 | -0.2 |
| Nemo-avg | +7.9 | +4.1 ∗ | +1.1 | +8.4 ∗ | -0.8 | +1.1 ∗ | |
| Qwen-2B | FW-avg | +2.7 | +3.2 | +0.3 | +3.5 | -0.4 | -0.0 |
| Nemo-avg | +8.7 ∗ | +14.6 ∗ | +0.5 | -0.7 | +0.0 | +1.1 ∗ | |
| Qwen-4B | FW-avg | +1.3 | -0.2 | +0.6 | +0.9 | -0.5 | +0.1 |
| Nemo-avg | +0.4 | +0.6 | +0.2 | -6.3 | -2.9 | -0.4 |
| Source corpus | BoolQ | RTE | OBQA | SciQ | TQA | RACE |
| HQ-DQA (8B, STEM Q&A) | +15.3 0.3 | +8.9 4.7 | +0.3 0.2 | +0.9 3.0 | +0.3 0.2 | +1.4 0.4 |
| HQ (26B, organic web) | +3.0 0.4 | +1.2 2.1 | +0.5 0.1 | -0.5 0.0 | +0.2 0.3 | -0.4 0.2 |
| FW-avg (reference) | +3.2 | +2.7 | +0.3 | +3.5 | -0.4 | -0.0 |
| Source | Phase 2 | BoolQ | RTE | SciQ |
| Mismatched Phase 2 (WikiText-103) | ||||
| 10M | WikiText-103 | +8.3 0.5 | +7.7 3.9 | -4.5 0.5 |
| 50M | WikiText-103 | +6.6 0.4 | +6.1 4.7 | -1.8 0.6 |
| 200M | WikiText-103 | +0.0 0.0 | +2.9 4.7 | -1.4 0.6 |
| Matched Phase 2 (HQ-DQA) | ||||
| 10M | HQ-DQA | +14.1 1.2 | +10.7 4.0 | -2.3 3.0 |
| Source | Gate mean | Frac. | ||
| tokens | WikiText-103 | BoolQ | WikiText-103 | BoolQ |
| 10M | 0.6 | 0.5 | 5.6% | 3.6% |
| 50M | 0.6 | 0.5 | 9.1% | 2.9% |
| 200M | 0.5 | 0.4 | 17.5% | 15.3% |
| Design | BoolQ | RTE | OBQA | SciQ | TQA | RACE | DQA-agg | Br-agg |
| Single-corpus references | ||||||||
| HQ-DQA ref. | +15.3 0.3 | +8.9 4.7 | +0.3 0.2 | +0.9 3.0 | +0.3 0.2 | +1.4 0.4 | +8.4 | +0.7 |
| FW-Edu ref. | +2.5 1.7 | +1.9 0.2 | -0.0 0.2 | +3.2 0.6 | -0.2 0.2 | +0.1 0.1 | +2.5 | -0.1 |
| Mixed or broadened memories | ||||||||
| 50/50 mix | +14.3 1.1 | +11.7 3.5 | +0.4 0.4 | -0.5 3.8 | +0.3 0.3 | +0.5 0.3 | +8.5 | +0.4 |
| Sequential HQ FW | +13.6 1.0 | +11.9 2.8 | -0.0 0.2 | -1.7 2.7 | +0.1 0.4 | +0.2 0.3 | +7.9 | +0.1 |
| Target Model | Dataset | Baseline | Transferred | (%) |
| Pythia-410M | LAMBADA | +0.0% | ||
| WikiText-103 | 22.1 | -2.4% | ||
| C4 | 24.8 | -0.3% | ||
| TinyLlama-1.1B | LAMBADA | 23.5 | -0.7% | |
| WikiText-103 | 10.1 | -5.0% | ||
| C4 | 11.6 | -0.2% |
| Reader corpus | Eval. corpus | No memory | Transfer | |
| WikiText | WikiText | 22.616 | 22.113 | -0.502 |
| WikiText | C4 | 24.726 | 24.663 | -0.062 |
| C4 | WikiText | 22.616 | 22.538 | -0.078 |
| C4 | C4 | 24.726 | 24.604 | -0.122 |
| Language | No-memory PPL | Transfer PPL | Rel. improvement |
| Chinese | 20.1248 | 2.251% | |
| Japanese | 15.3555 | 1.023% |
| Condition | Tokens | Train time | GPU-hours | Peak inf. memory | Prefill |
| No memory | 0 | 0 | 0 | GiB | ms |
| Phase-1 source artifact | 4.096M | 0.653 h | 3.154 | – | – |
| Fresh Mistral memory | 19.968M | 3.871 h | 15.760 | – | – |
| Matched FFN | 19.968M | 3.680 h | 14.929 | GiB | ms |
| Transfer, incl. Phase 1 | 4.096M+19.968M | 4.451 h | 18.623 | GiB | ms |