Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.
Figures & tables
Figure 1: Overview of the GSM architecture. The encoder progressively refines historical information, while the decoder layers share a single fixed-size KV memory.
Task
0.5B
1.5B
8B-A1.5B
Baseline
GSM
Baseline
GSM
Baseline
GSM
ARC-Easy † Clark et al. (2018)
51.85
51.64
59.09
58.59
72.90
71.55
ARC-Challenge †
28.75
26.88
33.45
33.36
36.69
36.69
PIQA † Bisk et al. (2020)
66.81
66.16
70.24
69.80
73.23
74.43
SciQ † Welbl et al. (2017)
74.10
73.10
81.90
82.40
89.80
89.80
BLiMP Warstadt et al. (2020)
77.01
79.34
80.18
80.79
81.66
82.24
Table 1: Downstream task performance (%). Metrics for tasks marked with † are specified in Appendix A.2 .
Answer Format
4K
8K
16K
32K
64K
Avg.
Direct Answer
98.33
100.00
98.33
100.00
98.33
99.00
Short CoT
100.00
100.00
100.00
100.00
100.00
100.00
Table 2: RULER retrieval accuracy of 1.5B GSM (%).
Model
s/step ↓
tokens/s ↑
GiB/GPU ↓
Baseline
2.578
50,833
43.54
GSM
2.462
53,231
44.06
Table 3: Training efficiency of the 1.5B models at a sequence length of 4K.
Retrieval Method
Standard Avg.
RULER Avg.
RULER 64K
Separate Aggregation Indexer
64.918
96.67
93.33
Shared Main Attention Top- K
64.765
94.33
93.33
Table 4: Shared retrieval ablation (0.5B, %).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Scale
Layers
FFN
Input Tokens
0.5B
24
Dense
10B
1.5B
28
Dense
15B
8B-A1.5B
28
MoE
18B
Appendix
Table 5: Main training configurations.
Evaluation Setting
Baseline
GSM
Main Training Held-Out Set
9.82440
9.94952
Full 64K After Adaptation
9.66497
9.66827
Appendix
Table 6: Supplementary FineWeb perplexity results for the 8B-A1.5B models.
Task
Direct Answer
Short CoT
niah_multikey_1
97.00
100.00
niah_single_1
100.00
100.00
niah_single_2
100.00
100.00
Appendix
Table 7: Per-task RULER accuracy of 1.5B GSM (%).
Length
Stage 1
Stage 2
Stage 3
4K
43.008
51.139
55.604
16K
28.478
34.225
43.645
32K
23.445
28.565
39.441
64K
19.023
24.676
36.170
Appendix
Table 8: Overlap between aggregation and main attention indices (%).
Modeling long-range dependencies remains a central challenge in natural language processing. Transformer architectures achieve strong performance via self-attention but scale quadratically (O(N2)) with sequence length, while State Space Models (SSMs) scale linearly (O(N)) but suffer from a selective recall bottleneck, struggling to retrieve precise information from compressed states. This creates a fundamental tradeoff between efficiency and perplexity. To tackle these challenges, we propose the \textit{Parallel Hybrid Architecture (PHA)}, which runs Gated State Spaces (GSS), Grouped Query Attention (GQA), and Feed-Forward Networks (FFNs) as independent parallel branches fused by a learnable mixing mechanism. Instead of forcing SSMs to approximate attention or serializing the two paradigms, PHA allows each branch to specialize: GSS captures global context, while attention performs selective retrieval, with FFN providing complementary processing. On WikiText-103, PHA achieves 16.51 PPL at 125M parameters, outperforming Hedgehog (16.70) and H3-125M (23.70). Scaling to 180M parameters yields 16.42 PPL, which gives comparable results with the pure attention baseline while delivering 24% higher throughput and up to 40% lower memory usage at long contexts. On OpenWebText, our 125M model achieves 19.72 PPL, outperforming standard Transformers (20.60) and GSS hybrid baselines (19.80). These results demonstrate that separating sequence modeling paradigms into parallel specialists enables Transformer-level perplexity with substantially improved efficiency for long-context language modeling.
Kuzey Torlak, Hüseyin Arda Arslan, Anıl Dervişoğlu +2
Kadıköy Anadolu High School · Politecnico di Torino · Istanbul Technical University +2
Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step state access dependent only on the number of selected slots rather than the total state size. Empirically, RAM-Net outperforms strong recurrent baselines on fine-grained long-range retrieval and achieves the lowest perplexity with competitive commonsense reasoning. It does so while accessing fewer state elements per step than all baselines, e.g., 8× fewer than Mamba2.
Modern language models are built primarily from Transformers, recurrent models, and their hybrid architectures. Transformers rely on token-level attention memories, while recurrent models such as state space models (SSMs) and linear attention maintain compact recurrent states. These architectures are typically instantiated separately or interleaved at the layer level, leaving open whether a shared memory representation can support both recurrent compression and attention-style retrieval. We study this question through the state space duality (SSD) view of Mamba-2, where the SSM state can be interpreted as a compressed associative key--value (KV) cache. We observe that Mamba-2 decodes token-conditioned values from this state but does not decode token-conditioned keys. Based on this observation, we propose DART (Decoded Attention over Recurrent sTates), which retains the chunk state contributions produced by the Mamba-2 chunked scan as chunk state memories, decodes token-conditioned keys and values from these memories, and performs state-memory attention (SMA) over the resulting KV pairs. The retrieved output is then combined with the native Mamba-2 output through a gated residual connection. DART supports practical training by reusing the Mamba-2 chunked scan and implementing SMA as a FlashAttention-style computation. Our analysis and experiments show that DART substantially reduces the length-dependent inference cache compared with a matched attention baseline (e.g., 75% savings when the chunk size is S=256 and the state size is N=128). Compared with Mamba-2, DART substantially improves associative recall and retrieval while preserving general language-modeling quality.
Yixiao Qian, Song Chen, Pengkai Wang +3
College of Control Science and Engineering, Zhejiang University · Department of Mathematics, National University of Singapore · Hong Kong Polytechnic University +1