Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce \emph{Hybrid Linear Attention} (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk's additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference. We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries. Project page: https://caesarhhh.github.io/hla/
Figures & tables
Figure 1 : Motivating observations for selective long-context retrieval. (a) In a pretrained GLA model, the cumulative recurrent retention of a state written at each historical position decreases with its distance to the final query position in a 4K-token prompt. The curve reports the median across layers and gate dimensions, with the shaded region denoting the interquartile range (IQR). (b) In pretrained GDN, the affine-state contribution of the chunk containing an early needle is progressively attenuated as it propagates through subsequent 512-token chunks. The curve reports the median normalized contribution over layer–head pairs, with IQR shading. (c) Chunk-level relevance can vary substantially across inputs. For a selected low-similarity LongBench pair, we report, at each layer, the minimum cross-sample cosine similarity over heads for softmax attention and HLA; MHLA remains at one because its chunk weights are input-independent. Together, these observations motivate query-dependent chunk-level attention for preserving and selectively accessing long-range information.
Figure 2 : Overview of HLA. Each completed chunk is represented by an affine transition and a set of self-attentively pooled routing representatives. For each query token, independent sigmoid gates interpolate historical transitions with the identity map. The gated transitions are composed in chronological order, followed by the causal prefix transition of the current chunk and the native GDN readout. The current chunk always has weight one.
Context
Method
Avg.
Single Needle
Multi-Key
Multi-Value/Query
VT
Word Extraction
QA
S1
S2
S3
MK1
MK2
MK3
MV
MQ
CWE
FWE
SQuAD
Hotpot
4K
GDN
24.63
99.00
66.00
46.00
26.60
0.00
0.00
15.80
15.60
2.56
24.06
5.20
11.00
8.40
4K
HLA (ours)
25.46
100.00
68.00
31.20
30.40
1.00
0.00
25.20
20.40
2.08
18.24
0.07
14.40
20.00
8K
GDN
15.02
56.60
29.80
20.60
20.40
0.00
0.00
8.50
7.65
13.76
17.28
5.33
6.60
8.80
8K
HLA (ours)
17.69
100.00
31.80
15.80
21.00
0.00
0.00
17.90
16.55
1.28
2.72
0.13
7.80
15.00
16K
GDN
7.98
23.80
9.40
2.40
9.60
0.00
0.00
3.90
1.90
22.64
6.08
9.80
5.20
9.00
Table 1 : From-scratch 1.3B GDN results on RULER from 4K to 32K. HLA and GDN are trained from scratch with the same 100B-token budget and a 4K training context. We report the macro average and the full task-level breakdown using 500 matched examples per task and context length. The better result within each context length is highlighted in bold.
Scale
Method
Avg.
Single Needle
Multi-Key
Multi-Value/Query
VT
Word Extraction
QA
S1
S2
S3
MK1
MK2
MK3
MV
MQ
CWE
FWE
SQuAD
Hotpot
0.8B
Native GDN
79.141
100.00
94.83
96.00
95.00
96.00
85.00
95.00
90.00
74.00
55.00
70.00
42.00
36.00
MHLA
80.182
100.00
95.33
95.60
95.30
96.50
86.00
95.80
89.70
76.20
58.20
71.50
44.50
37.74
HLA (ours)
83.115
100.00
96.83
95.00
96.00
98.00
90.00
97.00
89.00
81.00
67.00
76.00
52.00
42.67
2B
Native GDN
92.092
100.00
100.00
100.00
99.20
100.00
99.80
99.70
98.55
94.00
90.90
90.40
69.25
55.40
MHLA
92.192
100.00
100.00
99.80
99.50
100.00
100.00
99.85
98.20
94.30
91.20
90.70
69.60
55.34
Table 2 : RULER results across Qwen3.5 model scales. We report the macro average over all 13 tasks together with the complete task-level breakdown. Each task contains 500 examples. The best result within each model scale is highlighted in bold.
Scale
Method
Overall
Difficulty
Context Length
Easy
Hard
Short
Medium
Long
0.8B
Native GDN
22.86
23.96
22.19
25.56
19.07
25.93
MHLA
23.26
23.96
22.83
26.11
19.53
25.93
HLA (ours)
28.43
27.60
28.94
25.00
27.44
36.11
2B
Native GDN
27.83
31.25
25.72
26.67
29.77
25.93
MHLA
27.44
31.77
24.76
26.67
29.77
24.07
Table 3 : LongBench-V2 results across Qwen3.5 model scales. We report overall accuracy (%) together with breakdowns by difficulty and original context-length category. The best result within each model scale is highlighted in bold.
Figure 3 : Ablation analyses of chunk size, pooling window size, and historical-memory scaling.
Prefill
Decode
Scale
GDN (ms)
HLA (ms)
Ratio
GDN (ms/tok)
HLA (ms/tok)
Ratio
0.8B
40.36
46.00
1.140 ×
5.90
7.01
1.188 ×
2B
53.89
59.17
1.098 ×
6.49
7.60
1.171 ×
4B
115.09
132.29
1.149 ×
10.30
11.96
1.161 ×
9B
159.21
178.97
1.124 ×
12.19
13.88
1.139 ×
Table 4 : Prefill and decoding latency of Native GDN and HLA across Qwen3.5 model scales. We use a single NVIDIA H200 NVL, BF16, batch size 1, a 4096-token prompt, and 128-token steady-state decoding. HLA uses the same inference configuration as the main experiments.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Prefill
Decode / token
Inference Cache
Chunk Weighting
Softmax Attention
O(T2d)
O(Td)
O(Td)
Query-dependent
Linear Attention
O(Td2)
O(d2)
O(d2)
Implicit
GDN
O(Td2)
O(d2)
O(d2)
Implicit
MHLA
O(Td2+N2d2)
O(d2+Nd2/L)
O(Nd2)
Fixed
HLA (ours)
O(Tpd+TNPd+TMd2)
O(pd+NPd+Md2)
O(N(d2+Pd)+Ld)
Query-dependent
Appendix
Table 5 : Per-layer, per-head inference complexity. T is the sequence length, L the chunk size, N=⌈T/L⌉ the number of chunks, p the pooling-window size, P=L/p the number of representatives per chunk, and M the number of accessed chunks including the current chunk. We assume dk=dv=dr=d .
Figure 4 : Aggregate long-range historical contributions. Each line connects the same input under native GDN and HLA. Left: fraction of normalized historical contribution assigned to chunks in the distant half of the context. Right: contribution-weighted mean source distance. Black diamonds denote averages over the unfiltered analysis set. HLA shifts historical contribution toward more distant context under both measures.
Figure 5 : RULER S1 retrieval performance as a function of evidence distance. Results are grouped by the distance between the target evidence and the final query. Curves report answer-retrieval scores, with shaded intervals indicating uncertainty within each distance bin. HLA degrades substantially more slowly as the evidence moves farther from the query, particularly at 16K and 32K.
Figure 6 : Retrieval performance across context lengths and relative evidence positions. Each cell reports the RULER S1 retrieval score for examples grouped by the evidence-to-query distance as a fraction of the prompt length. Left: native GDN. Right: HLA. HLA retains substantially higher retrieval performance for distant evidence as the context length increases.
Figure 7 : Layer-wise contribution from distant history. From left to right, the panels correspond to 4K, 8K, 16K, and 32K. Each curve reports the fraction of normalized historical contribution assigned to the distant half of the context, averaged over paired examples. Shaded regions denote pointwise 95% bootstrap intervals. HLA exhibits substantially stronger distant-history contributions in several middle and upper layers, with the difference becoming more pronounced at longer contexts.
Figure 8 : Layer-wise contribution-weighted source distance. From left to right, the panels correspond to 4K, 8K, 16K, and 32K. The metric measures the average distance of historical sources, weighted by their normalized effective contributions. HLA increases the contribution distance in several middle and upper layers, particularly at 16K and 32K.
Figure 9 : Layer-wise contribution assigned to answer-evidence chunks. From left to right, the panels correspond to 4K, 8K, 16K, and 32K. We report the normalized historical contribution assigned to the chunk(s) containing the target evidence. HLA assigns greater relative contribution to evidence-containing chunks in many middle and upper layers, especially at longer context lengths.
Figure 10 : Historical-contribution maps at 4K and 8K context lengths. Each panel compares native GDN (left) and HLA (right) on the same RULER S1 examples. Rows correspond to different inputs and the vertical marker denotes the target-evidence position. HLA generally distributes non-negligible contribution over a broader range of historical positions.
Figure 11 : Historical-contribution maps at 16K and 32K context lengths. The visualization follows Figure 10 . At longer contexts, native GDN increasingly concentrates its historical contribution near the most recent positions, whereas HLA retains visible contributions from substantially more distant history, including regions around the target evidence in several examples. These visualizations are qualitative and do not constitute token-level causal attribution.
Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued pretraining on long-context data can help but is expensive due to the quadratic scaling of Attention. We observe that most tokens do not require (Global) Attention over the entire sequence and can rely on local context. Based on this, we propose L2A (Learning To Attend), a layer that enables conditional (token-wise) long-range memory access by deciding when to invoke global attention. We evaluate L2A on Qwen 2.5 and Qwen 3 models, extending their effective context length from 32K to 128K tokens. L2A matches the performance of standard long-context training to within 3% while skipping Global Attention for ∼80% of tokens, outperforming prior baselines. We also design custom Triton kernels to efficiently implement this token-wise conditional Attention on GPUs, achieving up to ∼2× improvements in training throughput and time-to-first-token over FlashAttention. Moreover, L2A enables post-training pruning of highly sparse Global Attention layers, reducing KV cache memory by up to 50% with negligible performance loss. Our code is released under Apache 2.0 at https://github.com/awslabs/hybrid-model-factory/tree/main/examples/research/L2A.
Hybrid large language models interleave full-attention layers with linear-attention layers to reduce the cost of long-context inference. This structure complicates prefix caching: full-attention key-value caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing hybrid prefix caching methods address this mismatch by storing recurrent-state checkpoints. As a result, token-level matches are directly usable only at positions aligned with stored checkpoints, constraining prefix reuse to a discrete set of boundaries. We present Tail-Replay, a prefix caching mechanism that enables unconstrained token-level prefix reuse in hybrid large language models. The key insight is that linear-attention mechanisms such as Gated DeltaNet can be viewed as a structured, lossy compression of the input prefix: gated recurrent updates progressively attenuate the contributions of earlier inputs. Consequently, the recurrent state of a matched prefix can be well approximated by replaying only a short, recent suffix of that prefix. Tail-Replay exploits this property by caching the exact full-attention key-value cache while omitting recurrent-state checkpoints. On a cache hit, it reconstructs the linear-attention states by replaying a short, recent suffix of the matched prefix. As a result, the reuse boundary is determined by the shared tokens rather than by recurrent-state checkpoints. We evaluate Tail-Replay on three Gated DeltaNet-based hybrid models using the LongBench and RULER benchmarks. With only a 5--10% replay budget, it retains 92.8--99.9% of full-prefill quality on LongBench and RULER. For serving efficiency, we evaluate time-to-first-token speedups across multiple matched-prefix lengths---8K, 16K, and 32K. The speedup grows with prefix length, reaching 9.1--14.3× over full prefill at 32K.
Yirui Liu, Ruoling Qi, Xuaner Wu +2
Institute of Artificial Intelligence, China Telecom (TeleAI) · Shanghai Jiao Tong University
Linear attention replaces the unbounded cache of softmax attention with a fixed-size recurrent state, reducing sequence mixing to linear time and decoding to constant memory. The hard part is not just what to forget, but how to edit this compressed memory without scrambling existing associations. Delta-rule models subtract the current read before writing a new value, and Kimi Delta Attention (KDA) sharpens forgetting with channel-wise decay. But the active edit still uses a single scalar gate to control two different things: how much old content to erase on the key side and how much new content to commit on the value side. We introduce Gated DeltaNet-2, which generalizes both Gated DeltaNet and KDA by inheriting adaptive forgetting and channel-wise decay while addressing their shared limitation, the scalar tie between erasing and writing. Gated Delta Rule-2 separates these roles with a channel-wise erase gate b_t and a channel-wise write gate w_t, reducing to KDA when both gates collapse to the same scalar and to Gated DeltaNet when the decay also collapses. We derive a fast-weight update view, a chunkwise WY algorithm with channel-wise decay absorbed into asymmetric erase factors, and a gate-aware backward pass that preserves efficient parallel training. At 1.3B parameters trained on 100B FineWeb-Edu tokens, Gated DeltaNet-2 achieves the strongest overall results among Mamba-2, Gated DeltaNet, KDA, and Mamba-3 variants across language modeling, commonsense reasoning, and retrieval. Its advantage is most pronounced on long-context RULER needle-in-a-haystack benchmarks, where it improves the evaluated multi-key retrieval setting and remains strong in both recurrent and hybrid settings. Code is available at https://github.com/NVlabs/GatedDeltaNet-2.