Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step state access dependent only on the number of selected slots rather than the total state size. Empirically, RAM-Net outperforms strong recurrent baselines on fine-grained long-range retrieval and achieves the lowest perplexity with competitive commonsense reasoning. It does so while accessing fewer state elements per step than all baselines, e.g., 8× fewer than Mamba2.
Figures & tables
Figure 1 : Comparison of attention architectures. (a) Full attention maintains KV cache and retrieves through competitive softmax over all cached keys, supporting high retrieval fidelity at the cost of memory that grows with sequence length t . (b) Linear attention compresses history into a fixed shared recurrent state and retrieves it through dense linear aggregation, thereby causing inter-token interference at retrieval. (c) RAM-Net keeps a fixed slot array and uses sparse address-based access: an Address Decoder maps kt and qt into sparse write and read addresses wt and rt , so only selected slots are updated or retrieved.
Figure 2 : Overview of RAM-Net. The Address Decoder maps the key kt and query qt to sparse weights over the slots through Product Softmax and Top- K truncation. Cyclic Address Positional Embedding (CAPE) then optionally shifts their slot indices by token position, producing the write and read weights wt and rt . Next, Gated Sparse Update writes vt into the selected slots, and the read output ot is normalized, gated, and projected. For visual clarity, the Address Decoder panel illustrates K=1 and ∣A∣=8 slots, with CAPE shifting the selected slot 5 to slot 4.
Figure 3 : Address Decoder. (a) The key kt or query qt is split into U sub-vectors of size dp , and each passes through its own softmax to give one digit distribution. (b) The Kronecker product of the digit distributions gives the address distribution over ∣A∣=dpU slots. (c) A Top- K tree search keeps the K most probable partial addresses at each level, returning the exact Top- K slots without materializing the ∣A∣ -dimensional distribution.
Model
Wiki. ppl ↓
LMB. acc ↑
MMLU acc ↑
ARC-e acc ↑
ARC-c acc ↑
OBQA accn↑
SciQ acc ↑
boolQ acc ↑
COPA acc ↑
PIQA acc ↑
Hella. acc ↑
Wino. acc ↑
Avg. acc ↑
Act. State 1 per step ↓
340M
Transformer++
26.81
33.1
23.1
57.7
24.7
32.0
82.9
55.9
66.0
65.8
33.3
52.2
49.4
100.7M
Gated DeltaNet
26.85
32.4
23.4
57.9
25.0
35.0
83.3
58.9
67.0
66.8
33.9
50.2
50.1
8.5M
Mamba2
26.23
33.9
23.1
59.2
24.2
33.8
83.4
59.6
71.0
67.7
33.9
50.4
50.6
12.9M
GLA
30.03
30.4
23.2
56.9
25.0
32.4
81.6
56.8
66.0
66.4
32.7
50.7
49.2
3.1M
GSA
29.17
32.7
23.0
59.3
25.8
35.8
83.0
57.4
66.0
66.4
33.3
49.8
50.0
3.1M
Table 1 : Zero-shot performance on language modeling and commonsense reasoning.
Figure 4 : MQAR accuracy versus state size. Active state in RAM-Net denotes the effective state accessed per step.
S-NIAH-1
S-NIAH-2
S-NIAH-3
Model
1K
2K
4K
8K
16K
1K
2K
4K
8K
16K
1K
2K
4K
8K
16K
Avg.
340M
Transformer++
99.7
97.7
81.0
0.0
0.0
100.0
99.7
74.7
0.0
0.0
81.7
62.7
31.3
0.0
0.0
48.6
Gated DeltaNet
99.3
99.3
98.7
88.0
44.0
99.7
96.0
37.3
20.3
10.3
46.0
26.7
2.3
0.7
0.3
51.3
Mamba2
89.3
91.7
77.0
24.7
6.7
97.3
72.7
12.0
11.7
4.0
20.7
10.3
5.3
1.3
0.7
35.0
GLA
68.7
23.7
5.7
0.3
0.3
64.7
20.3
6.3
3.3
1.0
5.0
0.3
0.0
0.0
0.0
13.3
Table 2 : Zero-shot performance comparison on S-NIAH benchmark. All models are trained with a 4K context window.
Model
FDA 512 ↑
FDA 1K ↑
SWDE 512 ↑
SWDE 1K ↑
NQ 512 ↑
NQ 1K ↑
SQuADv2 full ↑
TriviaQA full ↑
DROP full ↑
Avg.
340M
Transformer++
50.5
37.8
44.2
37.8
23.5
24.8
35.3
44.7
17.0
35.1
Gated DeltaNet
22.1
15.1
37.8
17.0
21.8
22.7
33.0
47.7
17.7
26.1
Mamba2
29.8
22.1
36.7
18.7
22.3
20.6
32.3
44.3
15.3
26.9
GLA
11.0
7.4
25.5
11.9
17.6
18.9
25.7
40.7
15.7
19.4
GSA
7.4
5.4
25.2
9.9
16.0
15.5
25.0
40.0
17.7
18.0
Table 3 : Zero-shot performance on recall-intensive tasks. FDA, SWDE, and NQ are evaluated at context lengths 512 and 1K; the remaining tasks use the full document.
Figure 5 : Decoding throughput on a single RTX 5090.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Setting
340M params.
1.3B params.
Data
Dataset
FineWeb-Edu [ 26 ]
Training budget
10B tokens
100B tokens
Tokenizer
Tokenizer
Llama 2 [ 44 ]
Vocabulary size
32k
Context window
4,096
Optimization
Optimizer
AdamW [ 24 ]
Appendix
Table 4 : Training configuration.
Figure 6 : Ablation study on the parameter order U in product softmax.
Variant
Wiki. ppl ↓
LMB. acc ↑
MMLU acc ↑
ARC-e acc ↑
ARC-c acc ↑
OBQA accn↑
SciQ acc ↑
boolQ acc ↑
COPA acc ↑
PIQA acc ↑
Hella. acc ↑
Wino. acc ↑
S-NIAH-1 acc ↑
S-NIAH-2 acc ↑
S-NIAH-3 acc ↑
RAM-Net Top-8 (default)
26.89
28.7
23.6
59.1
24.7
34.4
82.5
57.4
67.0
66.9
33.4
54.1
99.7
83.7
18.3
– Top-16
26.23
29.7
23.5
59.4
26.1
33.0
82.8
58.6
69.0
67.2
33.8
51.4
100.0
88.0
25.7
– Top-32
25.87
30.9
23.8
58.9
26.1
32.2
82.8
61.6
70.0
67.2
33.7
52.7
100.0
94.0
44.7
Ablations
– write 4, read 12
27.98
27.3
23.7
58.8
26.5
33.4
81.4
60.2
68.0
66.2
33.1
51.9
98.0
45.3
17.0
– write 12, read 4
26.59
29.3
23.1
59.7
25.8
33.8
82.8
62.0
70.0
66.3
33.6
51.1
100.0
48.7
7.7
Appendix
Table 5 : Ablation study of RAM-Net design choices. The default setting uses ∣A∣=45 slots, Top- K=8 , and the CAPE setting in Appendix A . Results are reported on language modeling, commonsense reasoning, and 4K S-NIAH retrieval tasks.
Figure 7 : Zoomed-in state access traces on the input “The original version of the transformer architecture was proposed in the 2017 paper ‘Attention Is All You Need’ by researchers at Google.” Each cell shows whether token t (vertical axis) issues a write (red) or read (green) event at slot address (horizontal axis); color intensity encodes the corresponding write or read strength. The left panel depicts a CAPE head, in which the cyclic shift produces a clear diagonal trace. The right panel depicts a NoPE head, in which addresses are determined purely by content.
Figure 8 : State access traces of four representative heads, showing write (red) and read (green) events at each slot address over tokens. CAPE heads (sub-figures (a) and (b)) produce diagonal traces, while NoPE heads (sub-figures (c) and (d)) show content-driven accesses.
Long-context recall in linear-time sequence models highlights a tradeoff in how they write to memory. State-based linear models, such as state-space models (SSMs) and linear Transformers, write densely, updating the entire state for each newly arrived token, which leads to interference and makes specific past tokens hard to recover. Sliding-window attention (SWA) exhibits the opposite behavior: it writes sparsely by storing explicit token representations, but only within a fixed window, so recall drops once the relevant token is evicted. Interpolating between these models, we introduce Raven, a linear-time sequence model that maintains a fixed set of memory slots and, at each step, decays and updates only a selected subset via learned, input-dependent routing. This lets Raven mitigate SWA's position-based overwriting and hard eviction while reducing interference from dense state updates in SSMs, thereby preserving long-range content much more effectively. Across recall-intensive benchmarks, Raven is competitive with or outperforms prior linear-time baselines, achieving strong long-context recall where both SWA and SSMs sharply degrade. It remains effective when extrapolating to context lengths as large as 16x its training length, with similar gains in hybrid architectures.
Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limited state size, linear attention models fall behind in long-context recall compared to softmax-attention-based transformer architectures. Increasing the state size of linear attention improves recall performance but at the cost of higher FLOPs. In this work, we introduce Sparse Delta Memory (SDM), an architecture that scales the hidden state of gated linear RNNs to orders of magnitude higher capacity using a sparse addressing scheme. SDM extends the Gated DeltaNet architecture by replacing the dense key-value outer product with sparse reads and writes to a large explicit memory. We show that, under an isoFLOP constraint and with an identical number of parameters, a higher state memory capacity significantly improves performance on in-context learning and long-context retrieval tasks. Moreover, by learning the initial state of the SDM memory and therefore using it as a parametric memory, we show that the model further improves on a wide range of common-knowledge and reasoning tasks.
Loïc Cabannes, Pierre-Emmanuel Mazaré, Gergely Szilvasy +6
Meta FAIR · Inria Paris & ENS-PSL University · University of Tübingen
Modern language models are built primarily from Transformers, recurrent models, and their hybrid architectures. Transformers rely on token-level attention memories, while recurrent models such as state space models (SSMs) and linear attention maintain compact recurrent states. These architectures are typically instantiated separately or interleaved at the layer level, leaving open whether a shared memory representation can support both recurrent compression and attention-style retrieval. We study this question through the state space duality (SSD) view of Mamba-2, where the SSM state can be interpreted as a compressed associative key--value (KV) cache. We observe that Mamba-2 decodes token-conditioned values from this state but does not decode token-conditioned keys. Based on this observation, we propose DART (Decoded Attention over Recurrent sTates), which retains the chunk state contributions produced by the Mamba-2 chunked scan as chunk state memories, decodes token-conditioned keys and values from these memories, and performs state-memory attention (SMA) over the resulting KV pairs. The retrieved output is then combined with the native Mamba-2 output through a gated residual connection. DART supports practical training by reusing the Mamba-2 chunked scan and implementing SMA as a FlashAttention-style computation. Our analysis and experiments show that DART substantially reduces the length-dependent inference cache compared with a matched attention baseline (e.g., 75% savings when the chunk size is S=256 and the state size is N=128). Compared with Mamba-2, DART substantially improves associative recall and retrieval while preserving general language-modeling quality.
Yixiao Qian, Song Chen, Pengkai Wang +3
College of Control Science and Engineering, Zhejiang University · Department of Mathematics, National University of Singapore · Hong Kong Polytechnic University +1