Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.
Figures & tables
Scale
Model
Wiki. ppl ↓
LMB. ppl ↓
LMB. acc ↑
PIQA ↑
Hella. ↑
Wino. ↑
Avg. score ↑
Avg. ppl ↓
340M
GLA
19.63
22.26
19.33
64.20
33.94
50.28
41.94
20.94
7B tok
MS-GLA {1,2}
19.25
20.99
19.68
65.18
34.89
49.17
42.23
20.12
MS-GLA {1,4}
17.43
21.69
19.89
63.87
35.49
51.07
42.58
19.56
MS-GLA {2,4}
21.66
31.47
16.15
65.18
34.17
52.57
42.01
26.56
MS-GLA {1,2,4}
18.33
19.59
20.69
65.83
35.01
49.88
42.85
18.96
MS-GLA {1,2,4,8}
18.80
20.18
20.34
65.34
34.53
49.01
42.31
19.49
Table 1: Language modeling perplexity and zero-shot downstream task performance of GLA and MS-GLA variants. Best results in each column are shown in bold .
Model
SWDE F1 ↑
FDA F1 ↑
SQuAD F1 ↑
Avg. F1 ↑
GLA
7.86
5.52
8.84
7.41
MS-GLA {1,2}
9.26
5.95
9.79
8.33
MS-GLA {1,4}
8.28
6.06
10.32
8.22
MS-GLA {2,4}
5.96
5.42
8.21
6.53
MS-GLA {1,2,4}
9.30
5.98
11.15
8.81
MS-GLA {1,2,4,8}
7.59
5.27
11.05
7.97
Table 2: Performance comparison of GLA and various MS-GLA configurations on recall-intensive benchmarks. F1 scores are reported for each task, with the highest values highlighted in bold .
Figure 1: Perplexity vs. Position-bucket on two datasets, PG19 (left) and SlimPajama (right) illustrating long-context generalization up to 30,720 tokens.
Region
MS-GLA {1,2}
MS-GLA {1,4}
MS-GLA {2,4}
MS-GLA {1,2,4}
MS-GLA {1,2,4,8}
First 10%
[72.93, 27.07]
[69.89, 30.11]
[49.81, 50.19]
[52.64, 18.41, 28.95]
[47.98, 18.86, 12.33, 20.83]
Middle 10%
[73.45, 26.55]
[70.00, 29.99]
[49.40, 50.60]
[52.91, 18.34, 28.76]
[48.35, 18.58, 12.16, 20.92]
Last 10%
[72.96, 27.03]
[69.47, 30.53]
[49.25, 50.75]
[52.57, 18.41, 29.03]
[48.04, 18.43, 12.53, 21.00]
Table 3: Average fusion routing weights (%) by scale, ordered finest to coarsest, at different positions within 32K-token evaluation sequences – roughly 15× the training context.
Model
TFLOPs
Throughput
Max mem. (GiB)
Train time
GLA
91.2
31,989
70.2
3d 1:49:25
MS-GLA {1,2}
78.9
27,619
86.7
2d 18:23:08
MS-GLA {1,2,4}
84.0
29,437
80.4
2d 22:48:41
MS-GLA {1,2,4,8}
77.8
27,201
89.2
2d 23:43:25
MS-GLA {2,4}
97.4
34,123
71.5
2d 9:11:04
Table 4: Training TFLOPs, throughput in tokens/sec, memory, and wall-clock time across GLA and MS-GLA variants, measured on the training runs described in Section 4.2 .