Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.
Figures & tables
Scale
Model
Wiki. ppl ↓
LMB. ppl ↓
LMB. acc ↑
PIQA ↑
Hella. ↑
Wino. ↑
Avg. score ↑
Avg. ppl ↓
340M
GLA
19.63
22.26
19.33
64.20
33.94
50.28
41.94
20.94
7B tok
MS-GLA {1,2}
19.25
20.99
19.68
65.18
34.89
49.17
42.23
20.12
MS-GLA {1,4}
17.43
21.69
19.89
63.87
35.49
51.07
42.58
19.56
MS-GLA {2,4}
21.66
31.47
16.15
65.18
34.17
52.57
42.01
26.56
MS-GLA {1,2,4}
18.33
19.59
20.69
65.83
35.01
49.88
42.85
18.96
MS-GLA {1,2,4,8}
18.80
20.18
20.34
65.34
34.53
49.01
42.31
19.49
Table 1: Language modeling perplexity and zero-shot downstream task performance of GLA and MS-GLA variants. Best results in each column are shown in bold .
Model
SWDE F1 ↑
FDA F1 ↑
SQuAD F1 ↑
Avg. F1 ↑
GLA
7.86
5.52
8.84
7.41
MS-GLA {1,2}
9.26
5.95
9.79
8.33
MS-GLA {1,4}
8.28
6.06
10.32
8.22
MS-GLA {2,4}
5.96
5.42
8.21
6.53
MS-GLA {1,2,4}
9.30
5.98
11.15
8.81
MS-GLA {1,2,4,8}
7.59
5.27
11.05
7.97
Table 2: Performance comparison of GLA and various MS-GLA configurations on recall-intensive benchmarks. F1 scores are reported for each task, with the highest values highlighted in bold .
Figure 1: Perplexity vs. Position-bucket on two datasets, PG19 (left) and SlimPajama (right) illustrating long-context generalization up to 30,720 tokens.
Region
MS-GLA {1,2}
MS-GLA {1,4}
MS-GLA {2,4}
MS-GLA {1,2,4}
MS-GLA {1,2,4,8}
First 10%
[72.93, 27.07]
[69.89, 30.11]
[49.81, 50.19]
[52.64, 18.41, 28.95]
[47.98, 18.86, 12.33, 20.83]
Middle 10%
[73.45, 26.55]
[70.00, 29.99]
[49.40, 50.60]
[52.91, 18.34, 28.76]
[48.35, 18.58, 12.16, 20.92]
Last 10%
[72.96, 27.03]
[69.47, 30.53]
[49.25, 50.75]
[52.57, 18.41, 29.03]
[48.04, 18.43, 12.53, 21.00]
Table 3: Average fusion routing weights (%) by scale, ordered finest to coarsest, at different positions within 32K-token evaluation sequences – roughly 15× the training context.
Model
TFLOPs
Throughput
Max mem. (GiB)
Train time
GLA
91.2
31,989
70.2
3d 1:49:25
MS-GLA {1,2}
78.9
27,619
86.7
2d 18:23:08
MS-GLA {1,2,4}
84.0
29,437
80.4
2d 22:48:41
MS-GLA {1,2,4,8}
77.8
27,201
89.2
2d 23:43:25
MS-GLA {2,4}
97.4
34,123
71.5
2d 9:11:04
Table 4: Training TFLOPs, throughput in tokens/sec, memory, and wall-clock time across GLA and MS-GLA variants, measured on the training runs described in Section 4.2 .
The scalability of Large Language Models (LLMs) to long contexts is fundamentally constrained by the quadratic complexity of standard attention, motivating the adoption of linear attention mechanisms with sub-quadratic cost. To improve representation capacity under long contexts, recent approaches organize memory in a multi-state manner. However, existing multi-state linear attention methods rely on fixed state merging policies that cannot adapt to dynamically varying token importance, irreversibly obscuring critical tokens and causing severe error accumulation over long sequences. To address this limitation, we propose DLA, a dynamic memory modeling framework for multi-state linear attention. DLA introduces (i) Information-Aware Dynamic State Merging, which adaptively determines state boundaries based on token-level information variation, preserving high-resolution representations around semantic transitions while aggressively summarizing stable regions, and (ii) Capacity-Bounded Memory Modeling, which maintains a fixed-size, chronologically ordered state cache by selectively merging adjacent low-information states to control memory growth with minimal information loss. We pre-train DLA on two different linear attention models and evaluate on 16 datasets across three categories. Experimental results demonstrate the superiority of DLA over state-of-the-art.
Xin Wang, Hui Shen, Boyuan Zheng +7
The Ohio State University · University of Michigan · ByteDance Seed
Real-world time series often exhibit irregular sampling and extended temporal horizons, requiring models to capture continuous-time dynamics across arbitrary intervals without prohibitive scaling costs. Discrete-time methods collapse variable time intervals into static positional steps; solver-dependent continuous-time models preserve temporal structure but rely on sequential integration, precluding parallelization; and solver-free approximations avoid this cost yet none couples observed time intervals with input-driven state modulation. We propose Liquid Gated Attention (LGA), a solver-free parallel temporal operator. By parameterizing an input-driven gating mechanism with observed time intervals, LGA introduces a continuous-time inductive bias and formulates hidden state evolution as a fast-weight associative memory, enabling parallel computation across the temporal dimension. Using matrix associativity in non-causal encoding and a prefix scan in causal encoding, LGA attains linear temporal complexity in sequence length in both modes. A sequence-level normalization bounds cumulative temporal decay for stable long-horizon optimization. Building on LGA, we instantiate LFormer, a modular backbone for continuous-time representation learning. Across six tasks and sixteen datasets spanning up to 17,984 steps, LFormer demonstrates long-range dependency modeling, fine-grained state tracking, and trajectory reconstruction from sparse and noisy observations, while delivering competitive performance against state-of-the-art discrete-time and continuous-time baselines with linear scaling efficiency.
Yiheng Jiang, Yuanbo Xu, Yongjian Yang
Mobile Intelligent Computing Laboratory, College of Computer Science and Technology, Jilin University, China
This paper introduces Exact Linear Attention (ELA), a mechanism that achieves linear computational complexity for Transformer attention by exploiting the exact decomposition property of kernel functions, thereby eliminating approximation error. We identify and address two key limitations of prior linear attention -- gradient explosion and token attention dilution -- by imposing kernel constraints that ensure non-negativity, discriminability, and geometric interpretability. Several kernel functions are proposed, including the Hadamard Exp Kernel, Summation Squared Euclidean Distance Kernel, and Subtraction Squared Euclidean Distance Kernel, each tailored for specific attention behaviors. Beyond the core attention formulation, the paper presents three engineering innovations: (1) a Hyper-Link structure that replaces traditional residual connections to mitigate gradient degradation; (2) a Memory Lobe module based on bidirectional linear attention, which captures "transformation flow" across layers to implement qualitative memory and an implicit reinforcement learning paradigm; and (3) a routing-score-based bias mechanism for Mixture-of-Experts (MoE) to improve interpretability and semantic alignment. Experimental results demonstrate that ELA achieves up to 6x faster decoding speed and 75% reduction in KV cache memory usage compared to full attention, while maintaining comparable or superior training performance. The proposed memory module accelerates convergence and enhances generalization. Furthermore, we extend the linear attention principle to vision models, yielding YOLO-LAT, which attains up to 4.3x GPU inference speedup and 7.9x parameter reduction with competitive detection accuracy. These results underline the broad applicability of exact linear attention for scaling Transformer models to ultra-long sequences and efficient visual tasks.