MemoryAthena: Adaptive Routing over Latent and Generated Memories
Organizations: ELLIS Institute of Finland · University of Turku · University of Technology Sydney · University of Science and Technology of China · Shanghai Jiao Tong University
Abstract
Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. MemoryAthena uses three pathways: direct Engram retrieval (E), generation from retrieved Engram cues (GE), and generation from causal backbone states without consulting the memory table (GH). Generated memory is conditionally useful: it can complement E in one context but interfere with it in another. MemoryAthena therefore treats E as an anchor and learns when a generated representation should intervene. With the backbone, memory, generators, and readers frozen, a lightweight causal routing head is trained from counterfactual future-token likelihood advantages of GE and GH relative to E. At inference time, an admitted candidate modifies the E residual through bounded interpolation, while rejection recovers the direct pathway exactly. On question answering, MemoryAthena raises the five-task average from 37.65 to 39.28 over the direct pathway of the same checkpoint, while the six-task general-NLP average increases from 76.73 to 79.13. The complete memory-side system contains approximately 201M parameters, excluding the frozen backbone. Further analyses show complementary strengths among E, GE, and GH across tasks and inputs. These results support generated memory as a selective correction to direct retrieval and highlight routing when, which, and how strongly to intervene as the central challenge.
Figures & tables
| Method / Inference rule | NQ | WebQA | TriviaQA | TruthfulQA | HotpotQA | Average |
|---|---|---|---|---|---|---|
| Reported baselines | ||||||
| Base (Vanilla Mistral) | 20.60 | 29.30 | 57.70 | 32.10 | 21.00 | 32.14 |
| RAG | 22.60 +2.00 | 24.90 -4.40 | 54.20 -3.50 | 35.50 +3.40 | 29.80 +8.80 | 33.40 (+3.9%) |
| kNN-LM | 21.10 +0.50 | 30.50 +1.20 | 57.80 +0.10 | 32.30 +0.20 | 21.20 +0.20 | 32.58 (+1.4%) |
| CPT | 12.20 -8.40 | 34.10 +4.80 | 61.20 +3.50 | 29.20 -2.90 | 16.00 -5.00 | 30.54 (-5.0%) |
| LoRA | 18.20 -2.40 | 34.50 +5.20 | 61.60 +3.90 | 30.90 -1.20 | 16.20 -4.80 | 32.28 (+0.4%) |
| Method | SST2 | MR | CR | RT | AGN | Yahoo | Average |
|---|---|---|---|---|---|---|---|
| Full-choice dCPMI baseline | |||||||
| Mistral-7B-v0.3 | 81.08 | 75.60 | 74.00 | 74.67 | 73.24 | 55.03 | 72.27 |
| Non-parametric methods (reported) | |||||||
| RAG | 87.20 +6.12 | 83.70 +8.10 | 71.55 -2.45 | 82.36 +7.69 | 75.64 +2.40 | 58.43 +3.40 | 76.48 (+5.8%) |
| kNN-LM | 82.15 +1.07 | 76.85 +1.25 | 61.70 -12.30 | 74.95 +0.28 | 76.13 +2.89 | 56.26 +1.23 | 71.34 (-1.3%) |
| Parametric methods (reported) | |||||||
| Condition | NQ | WebQA | TriviaQA | TruthfulQA | HotpotQA | Average |
|---|---|---|---|---|---|---|
| Base (Vanilla Mistral, reported) | 20.60 | 29.30 | 57.70 | 32.10 | 21.00 | 32.14 |
| Ours, pretrained memory | 33.02 +12.42 | 34.60 +5.30 | 70.68 +12.98 | 31.74 -0.36 | 26.34 +5.34 | 39.28 (+22.2%) |
| Ours, from scratch | 33.79 +13.19 | 34.26 +4.96 | 72.97 +15.27 | 32.09 -0.01 | 27.71 +6.71 | 40.16 (+25.0%) |
| Interface and addressing controls | ||||||
| No gate | 27.49 +6.89 | 30.88 +1.58 | 55.15 -2.55 | 32.15 +0.05 | 21.10 +0.10 | 33.35 (+3.8%) |
| Permuted Engram keys | 32.72 +12.12 | 31.24 +1.94 | 70.06 +12.36 | 32.62 +0.52 | 26.24 +5.24 | 38.58 (+20.0%) |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Trainable components | Budget |
|---|---|---|
| Memory learning | Engram memory table and source-side adaptor | 20M |
| Memory-interface adaptation | GE/GH generators and target-side readers | 20M |
| Router training | E-relative advantage and confidence heads | 20M |
| Evaluation | None; all components are frozen | – |
| Setting | Value |
|---|---|
| Packed sequence length | 2,048 |
| Training steps | 9,765 |
| Processed input positions | 19,998,720 |
| Validation budget | 2,000,000 positions |
| Trainable router parameters | 534,924 |
| Router hidden width | 64 |
| Component | Configuration | Value |
| Memory interface | ||
| Injection layers | Target backbone layers | |
| Memory dimension | Retrieved / latent memory width | 512 |
| Memory table | Frozen Engram parameters | 33,554,432 |
| Reader branches | Parallel branches per injection layer | 4 |
| Direct E reader | Parameters per layer / two layers | 10,506,244 / 21,012,488 |
| Condition | MC1 | MC2 | MC3 | Mean |
|---|---|---|---|---|
| Standalone Engram | 27.42 | 44.32 | 22.69 | 31.47 |
| Same-checkpoint E | 26.93 | 44.18 | 22.64 | 31.25 |
| E/GE learned pair | 27.54 | 43.40 | 22.72 | 31.22 |
| E/GH learned pair | 27.78 | 42.90 | 22.67 | 31.12 |
| MemoryAthena | 28.15 | 44.04 | 23.02 | 31.74 |
| Mistral Llama | 27.78 | 41.57 | 21.87 | 30.41 |
| Task | E | GE | GH |
|---|---|---|---|
| NQ | 2,854 (79.08) | 475 (13.16) | 280 (7.76) |
| WebQA | 1,428 (70.28) | 336 (16.54) | 268 (13.19) |
| TriviaQA | 15,764 (87.85) | 1,345 (7.50) | 835 (4.65) |
| TruthfulQA | 361 (44.19) | 254 (31.09) | 202 (24.72) |
| HotpotQA | 6,150 (83.05) | 762 (10.29) | 493 (6.66) |
| Threshold | Accuracy (%) |
|---|---|
| 0.00 | 45.8967 |
| 0.05 | 46.1150 |
| 0.10 | 46.2483 |
| 0.20 | 47.0283 |
| 0.30 | 48.3583 |
| 0.90 | 57.3600 |
| Validation corpus | E | GE | GH |
|---|---|---|---|
| Wikipedia-2021, QA checkpoint | 23.59 | 7.56 | 68.85 |
| General mixture, NLP checkpoint | 62.85 | 4.49 | 32.67 |
| Nemotron-CC-Code, coding checkpoint | 64.80 | 5.11 | 30.09 |
| Condition | SST2 | MR | CR | RT | AGN | Yahoo | Mean |
|---|---|---|---|---|---|---|---|
| Target E | 50.92 | 44.85 | 50.00 | 50.00 | 23.62 | 10.00 | 38.23 |
| Target router | 49.08 | 47.95 | 58.35 | 52.44 | 25.05 | 10.00 | 40.48 |
| Method | QA | Summarization | Average |
|---|---|---|---|
| Reported baselines | |||
| Mistral-7B-v0.3 | 53.99 | 50.27 | 52.13 |
| CPT | 46.49 -7.50 | 47.39 -2.88 | 46.94 |
| LoRA | 50.02 -3.97 | 50.38 +0.11 | 50.20 |
| RAG | 65.09 +11.10 | – | – |
| MLP Memory | 64.07 +10.08 | 52.41 +2.14 | 58.24 |
| Scale | Backbone | Engram | Gen. | Readers | Router | Memory-side |
|---|---|---|---|---|---|---|
| Small | 124M | 33.554M | 256/2/4 | 16 | 64 | 37.573M |
| Medium | 345M | 93.716M | 428/2/6 | 24 | 104 | 104.008M |
| Large | 774M | 209.715M | 640/3/8 | 40 | 160 | 238.212M |
| XL | 1.5B | 405.537M | 896/4/12 | 56 | 224 | 472.912M |
| Corpus | Scale | No memory | Memory | Router |
|---|---|---|---|---|
| WikiText | Small | 30.841 | 24.033 | 23.372 |
| Medium | 22.569 | 17.797 | 17.291 | |
| Large | 19.342 | 14.955 | 14.545 | |
| General mixture | Small | 36.752 | 32.300 | 31.772 |
| Medium | 27.958 | 24.445 | 24.181 |
| Training budget | Memory | Router |
|---|---|---|
| 10M | 20.714 | 20.222 |
| 30M | 20.369 | 19.889 |
| 100M | 19.783 | 19.462 |
| Setting | Stage | Trainable params. | Time (h) |
|---|---|---|---|
| QA | Reader / generator | 167.386M | 31.70 |
| Router | 0.535M | 11.19 | |
| General NLP | Reader / generator | 167.365M | 16.01 |
| Router | 0.535M | 10.58 | |
| Coding | Reader / generator | 167.365M | 16.41 |
| Router | 0.535M | 10.94 |
| E | GE | GH | Router |
| 9.07260 | 8.79091 | 8.78358 | 8.83063 |