RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State
Organizations: The Chinese University of Hong Kong
Abstract
Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step state access dependent only on the number of selected slots rather than the total state size. Empirically, RAM-Net outperforms strong recurrent baselines on fine-grained long-range retrieval and achieves the lowest perplexity with competitive commonsense reasoning. It does so while accessing fewer state elements per step than all baselines, e.g., fewer than Mamba2.
Figures & tables
| Model | Wiki. ppl | LMB. acc | MMLU acc | ARC-e acc | ARC-c acc | OBQA | SciQ acc | boolQ acc | COPA acc | PIQA acc | Hella. acc | Wino. acc | Avg. acc | Act. State 1 per step |
| 340M | ||||||||||||||
| Transformer++ | 26.81 | 33.1 | 23.1 | 57.7 | 24.7 | 32.0 | 82.9 | 55.9 | 66.0 | 65.8 | 33.3 | 52.2 | 49.4 | 100.7M |
| Gated DeltaNet | 26.85 | 32.4 | 23.4 | 57.9 | 25.0 | 35.0 | 83.3 | 58.9 | 67.0 | 66.8 | 33.9 | 50.2 | 50.1 | 8.5M |
| Mamba2 | 26.23 | 33.9 | 23.1 | 59.2 | 24.2 | 33.8 | 83.4 | 59.6 | 71.0 | 67.7 | 33.9 | 50.4 | 50.6 | 12.9M |
| GLA | 30.03 | 30.4 | 23.2 | 56.9 | 25.0 | 32.4 | 81.6 | 56.8 | 66.0 | 66.4 | 32.7 | 50.7 | 49.2 | 3.1M |
| GSA | 29.17 | 32.7 | 23.0 | 59.3 | 25.8 | 35.8 | 83.0 | 57.4 | 66.0 | 66.4 | 33.3 | 49.8 | 50.0 | 3.1M |
| S-NIAH-1 | S-NIAH-2 | S-NIAH-3 | ||||||||||||||
| Model | 1K | 2K | 4K | 8K | 16K | 1K | 2K | 4K | 8K | 16K | 1K | 2K | 4K | 8K | 16K | Avg. |
| 340M | ||||||||||||||||
| Transformer++ | 99.7 | 97.7 | 81.0 | 0.0 | 0.0 | 100.0 | 99.7 | 74.7 | 0.0 | 0.0 | 81.7 | 62.7 | 31.3 | 0.0 | 0.0 | 48.6 |
| Gated DeltaNet | 99.3 | 99.3 | 98.7 | 88.0 | 44.0 | 99.7 | 96.0 | 37.3 | 20.3 | 10.3 | 46.0 | 26.7 | 2.3 | 0.7 | 0.3 | 51.3 |
| Mamba2 | 89.3 | 91.7 | 77.0 | 24.7 | 6.7 | 97.3 | 72.7 | 12.0 | 11.7 | 4.0 | 20.7 | 10.3 | 5.3 | 1.3 | 0.7 | 35.0 |
| GLA | 68.7 | 23.7 | 5.7 | 0.3 | 0.3 | 64.7 | 20.3 | 6.3 | 3.3 | 1.0 | 5.0 | 0.3 | 0.0 | 0.0 | 0.0 | 13.3 |
| Model | FDA 512 | FDA 1K | SWDE 512 | SWDE 1K | NQ 512 | NQ 1K | SQuADv2 full | TriviaQA full | DROP full | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|
| 340M | ||||||||||
| Transformer++ | 50.5 | 37.8 | 44.2 | 37.8 | 23.5 | 24.8 | 35.3 | 44.7 | 17.0 | 35.1 |
| Gated DeltaNet | 22.1 | 15.1 | 37.8 | 17.0 | 21.8 | 22.7 | 33.0 | 47.7 | 17.7 | 26.1 |
| Mamba2 | 29.8 | 22.1 | 36.7 | 18.7 | 22.3 | 20.6 | 32.3 | 44.3 | 15.3 | 26.9 |
| GLA | 11.0 | 7.4 | 25.5 | 11.9 | 17.6 | 18.9 | 25.7 | 40.7 | 15.7 | 19.4 |
| GSA | 7.4 | 5.4 | 25.2 | 9.9 | 16.0 | 15.5 | 25.0 | 40.0 | 17.7 | 18.0 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Setting | 340M params. | 1.3B params. |
|---|---|---|---|
| Data | Dataset | FineWeb-Edu [ 26 ] | |
| Training budget | 10B tokens | 100B tokens | |
| Tokenizer | Tokenizer | Llama 2 [ 44 ] | |
| Vocabulary size | 32k | ||
| Context window | 4,096 | ||
| Optimization | Optimizer | AdamW [ 24 ] | |
| Variant | Wiki. ppl | LMB. acc | MMLU acc | ARC-e acc | ARC-c acc | OBQA | SciQ acc | boolQ acc | COPA acc | PIQA acc | Hella. acc | Wino. acc | S-NIAH-1 acc | S-NIAH-2 acc | S-NIAH-3 acc |
| RAM-Net Top-8 (default) | 26.89 | 28.7 | 23.6 | 59.1 | 24.7 | 34.4 | 82.5 | 57.4 | 67.0 | 66.9 | 33.4 | 54.1 | 99.7 | 83.7 | 18.3 |
| – Top-16 | 26.23 | 29.7 | 23.5 | 59.4 | 26.1 | 33.0 | 82.8 | 58.6 | 69.0 | 67.2 | 33.8 | 51.4 | 100.0 | 88.0 | 25.7 |
| – Top-32 | 25.87 | 30.9 | 23.8 | 58.9 | 26.1 | 32.2 | 82.8 | 61.6 | 70.0 | 67.2 | 33.7 | 52.7 | 100.0 | 94.0 | 44.7 |
| Ablations | |||||||||||||||
| – write 4, read 12 | 27.98 | 27.3 | 23.7 | 58.8 | 26.5 | 33.4 | 81.4 | 60.2 | 68.0 | 66.2 | 33.1 | 51.9 | 98.0 | 45.3 | 17.0 |
| – write 12, read 4 | 26.59 | 29.3 | 23.1 | 59.7 | 25.8 | 33.8 | 82.8 | 62.0 | 70.0 | 66.3 | 33.6 | 51.1 | 100.0 | 48.7 | 7.7 |