Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.
Figures & tables
Figure 1: Overview of the GSM architecture. The encoder progressively refines historical information, while the decoder layers share a single fixed-size KV memory.
Task
0.5B
1.5B
8B-A1.5B
Baseline
GSM
Baseline
GSM
Baseline
GSM
ARC-Easy † Clark et al. (2018)
51.85
51.64
59.09
58.59
72.90
71.55
ARC-Challenge †
28.75
26.88
33.45
33.36
36.69
36.69
PIQA † Bisk et al. (2020)
66.81
66.16
70.24
69.80
73.23
74.43
SciQ † Welbl et al. (2017)
74.10
73.10
81.90
82.40
89.80
89.80
BLiMP Warstadt et al. (2020)
77.01
79.34
80.18
80.79
81.66
82.24
Table 1: Downstream task performance (%). Metrics for tasks marked with † are specified in Appendix A.2 .
Answer Format
4K
8K
16K
32K
64K
Avg.
Direct Answer
98.33
100.00
98.33
100.00
98.33
99.00
Short CoT
100.00
100.00
100.00
100.00
100.00
100.00
Table 2: RULER retrieval accuracy of 1.5B GSM (%).
Model
s/step ↓
tokens/s ↑
GiB/GPU ↓
Baseline
2.578
50,833
43.54
GSM
2.462
53,231
44.06
Table 3: Training efficiency of the 1.5B models at a sequence length of 4K.
Retrieval Method
Standard Avg.
RULER Avg.
RULER 64K
Separate Aggregation Indexer
64.918
96.67
93.33
Shared Main Attention Top- K
64.765
94.33
93.33
Table 4: Shared retrieval ablation (0.5B, %).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Scale
Layers
FFN
Input Tokens
0.5B
24
Dense
10B
1.5B
28
Dense
15B
8B-A1.5B
28
MoE
18B
Appendix
Table 5: Main training configurations.
Evaluation Setting
Baseline
GSM
Main Training Held-Out Set
9.82440
9.94952
Full 64K After Adaptation
9.66497
9.66827
Appendix
Table 6: Supplementary FineWeb perplexity results for the 8B-A1.5B models.
Task
Direct Answer
Short CoT
niah_multikey_1
97.00
100.00
niah_single_1
100.00
100.00
niah_single_2
100.00
100.00
Appendix
Table 7: Per-task RULER accuracy of 1.5B GSM (%).
Length
Stage 1
Stage 2
Stage 3
4K
43.008
51.139
55.604
16K
28.478
34.225
43.645
32K
23.445
28.565
39.441
64K
19.023
24.676
36.170
Appendix
Table 8: Overlap between aggregation and main attention indices (%).
College of Control Science and Engineering, Zhejiang University · Department of Mathematics, National University of Singapore · Hong Kong Polytechnic University +1