The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16× training-free length extrapolation while maintaining 100% accuracy on NIAH-SK1 in 64k context length.
Figures & tables
Figure 1 : Comparison of training loss, perplexity (PPL), and LongPPL among different layer-wise hybrid models after 4k context length pretraining. SWA/GLA/GDN-NoPE hybrids achieve stronger long-context language modeling.
Figure 2 : Comparison of long-context performance under RULER and BABILong benchmarks among different layer-wise hybrid models after 4k context length pretraining. SWA/GLA/GDN-NoPE hybrids achieve both stronger long-context downstream performance within and beyond the training context length.
Figure 3 : Comparison of training loss, perplexity (PPL), and LongPPL among different head-wise hybrid models after 4k context length pretraining. SWA/GLA/GDN-NoPE hybrids achieve stronger long-context language modeling.
Figure 4 : Comparison of long-context performance under RULER and BABILong benchmarks among different head-wise hybrid models after 4k context length pretraining. SWA/GLA/GDN-NoPE hybrids achieve both stronger long-context downstream performance within and beyond the training context length.
Figure 5 : Comparison of long-context performance between RoPE-only and RoPE-NoPE hybrid full-attention models. Both layer-wise and head-wise RoPE-NoPE hybrids have better long-context performance and can achieve effective length extrapolation directly or under log-scaled NoPE extrapolation.
Figure 6 : Comparison of long-context performance under RULER and BABILong benchmarks among different layer-wise hybrid models after 32k context length continual pretraining.
Figure 7 : Comparison of long-context performance under RULER and BABILong benchmarks among different head-wise hybrid models after 32k context length continual pretraining.
Figure 8 : Comparison of the long-context performance of the SWA/GLA/GDN-NoPE layer-wise hybrid models under log-scale NoPE extrapolation after short-context pretraining (left) and after long-context continual pretraining (right). The results of SWA hybrids exhibit a remarkable seesaw effect, as detailed in Takeaway 2.
Figure 9 : Comparison of the long-context performance of the SWA/GLA/GDN-NoPE head-wise hybrid models under log-scale NoPE extrapolation after short-context pretraining (left) and after long-context continual pretraining (right). The results of SWA hybrids exhibit a remarkable seesaw effect, as detailed in Takeaway 2.
Figure 10 : Comparison of training loss, RULER, and BABILong among different hybrid ratios in head-wise hybrid models.
Figure 11 : Results of noise experiments on RULER and BABILong benchmark. Adding Gaussian noise to RoPE full attention or SWA/GLA/GDN degrades long-context performance more severely than the same perturbation applied to NoPE full attention in the corresponding hybrid models.
Figure 12 : The results of attention analysis within 1:1 RoPE-NoPE-HH and among RoPE-NoPE-HH with different hybrid ratio. NoPE full attention shows a stronger tendency for coarse global information localization, given the higher attention score hit rate and larger attention entropy, compared with RoPE full attention in RoPE-NoPE hybrids.
Figure 13 : The results of attention analysis within 1:1 RoPE-NoPE-LH and among RoPE-NoPE-LH with different hybrid ratios. The tidal effect of hybrid position is still applicable in RoPE-NoPE layer-wise hybrid models.
Figure 14
Figure 14 : Perplexity curves of different hybrid models under direct extrapolation, log-scale NoPE extrapolation (Log), and long-context continual pretraining (LC-CPT). The perplexity curves of SWA-NoPE-LH show a remarkable increasing trend as the context length grows, different from other hybrid models.
Figure 15 : Comparison of the change rates of the WQ and WK matrices in attention modules across different hybrid models before and after long-context training. In layer-wise hybrids, 1 NoPE layer is inserted every 3 RoPE/SWA/GLA/GDN layers. In head-wise hybrids, each layer contains the WQ and WK matrices of NoPE as well as those of the other RoPE/SWA/GLA/GDN modules. Remarkably, in SWA-NoPE hybrids, the SWA modules receive anomalous emphasis during long-context training. However, their limited window size cannot support modeling longer contexts.
Figure 17
Figure 18 : Tendency of long-context perplexity and performance as the window size of SWA-NoPE-LH increases in long-context pretraining. Larger window sizes show better context extension, implying short-window weariness.
Figure 19 : Tendency of long-context perplexity and performance as the window size of SWA-NoPE-LH increases in short-context pretraining. Shorter window sizes show better length extrapolation, implying long-window laziness.
Figure 20 : Comparison of long-context performance of SWA-NoPE-LH after long-context continual pretraining enhanced by LongCE and extended window size. Both clearly weaken the seesaw effect and also contribute to SWA-RoPE-LH.
Figure 21 : Attention analysis results among 3:1 RoPE-NoPE, SWA-NoPE, GLA-NoPE, and GDN-NoPE hybrid models, with extended attention entropy and hit rate. In the entropy-hit-rate diagram, NoPE attention still shows a stronger tendency to aggregate global information coarsely, with a higher attention hit rate and larger attention entropy. In contrast, SWA, GLA, and GDN show a similar distribution pattern to RoPE attention.
Figure 22 : Attention analysis results among GLA-NoPE-HH and GDN-NoPE-HH with different hybrid ratios. The tidal effect of hybrid position is still applicable in GLA-NoPE and GDN-NoPE head-wise hybrid models.
Figure 23 : Comparison of length extrapolation in GLA-NoPE hybrid models. Applying a sliding window for GLA and a log scale for NoPE attention achieves the best length extrapolation performance.
Figure 24 : Comparison of length extrapolation in GDN-NoPE hybrid models. Still, applying a sliding window for GDN and a log scale for NoPE attention achieves the best length extrapolation performance in most cases.
Figure 25
Figure 25 : Comparison of EME in RoPE-NoPE, GLA-NoPE, and GDN-NoPE hybrids as well as SWA-NoPE hybrids.
NIAH-SK1
NIAH-SK2
NIAH-SK3
4k
8k
16k
32k
64k
4k
8k
16k
32k
64k
4k
8k
16k
32k
64k
776M Short
GLA-NoPE-LH
99.0
95.0
80.0
0.0
0.0
99.0
69.0
0.0
0.0
0.0
90.0
88.0
0.0
0.0
0.0
+ Log
99.0
99.0
98.0
0.0
0.0
99.0
68.0
1.0
0.0
0.0
90.0
91.0
0.0
0.0
0.0
+ SWLA
99.0
100.0
100.0
97.0
97.0
99.0
74.0
5.0
0.0
0.0
90.0
56.0
21.0
8.0
3.0
+ SWLA & Log (EME)
99.0
100.0
100.0
96.0
100.0
99.0
92.0
87.0
72.0
59.0
90.0
79.0
75.0
76.0
64.0
Table 1: Extrapolation comparison of the 776M GLA/GDN-NoPE hybrid model on the NIAH-SK1, SK2, and SK3 tasks. Applying a sliding-window linear attention and a log scale for NoPE attention achieves the best length extrapolation.
Figure 26 : A brief comparison between the traditional full attention model and hybrid models on inference efficiency.
Figure 27 : A brief comparison between the traditional full attention model and hybrid models on training efficiency.
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
376M
776M
1B
3B
Hidden Size
1024
1536
2048
3072
Intermediate Size
3584
5376
7168
10752
Num Layer
8
12
16
24
Num Attn Head
8
12
16
24
Num KV Head
8
12
16
24
Vocab Size
128256
128256
128256
128256
Appendix
Table 2: The hyperparameters of different model sizes.
Figure 28 : Comparison of the LongBench performance of the SWA/GLA/GDN-NoPE hybrid models under log-scale NoPE extrapolation (Short-Log) and context extension (Long). The Seesaw Effect still exists.
Figure 29 : Comparison of LongBench performance in RoPE-NoPE hybrid models. Applying a sliding-window RoPE attention and a log-scale NoPE attention achieves the best length extrapolation. The Matthew Effect still works.
Figure 30 : Comparison of LongBench performance in GLA-NoPE hybrid models. Applying a sliding-window GLA and a log-scale NoPE attention achieves the best length extrapolation. The Matthew Effect still works.
Figure 31 : Comparison of LongBench performance in GDN-NoPE hybrid models. Applying a sliding-window GDN and a log-scale NoPE attention achieves the best length extrapolation. The Matthew Effect still works.
Figure 32 : Comparison of long-context training performance between SWA-NoPE-LH, GLA-NoPE-LH, and GDN-NoPE-LH under different learning rates, where 0 learning rate means the performance before long-context continual pretraining.
Figure 33 : Comparison of the long-context performance of the SWA/GLA/GDN-NoPE layer-wise hybrid models in YOCO-like layout under log-scale NoPE extrapolation and context extension. The Seesaw Effect still exists.
Figure 34 : Performance of enhanced SWA-NoPE-LH context extension and RoPE/GLA/GDN-NoPE-LH extrapolation based on the Matthew Effect in layer-wise hybrid models in YOCO-like layout.
Figure 35 : The entropy-hit-rate diagram of RoPE-based full attention models and SWA/GLA/GDN-RoPE hybrid models.
Figure 36 : The entropy-hit-rate diagram of NoPE/RoPE-based full attention models and RoPE-NoPE-HH models.
TQA
LMD
PIQA
Hella
Wino
ARC-e
ARC-c
GPQA
SIQA
OBQA
SG
MMLU
Avg.
376M Short
RoPE
36.8
40.2
65.9
33.8
52.2
39.7
26.4
26.3
39.2
27.4
42.1
24.6
37.9
RoPE-NoPE-LH
36.4
41.7
65.9
33.8
51.1
39.2
27.5
28.8
40.3
25.8
45.2
26.9
38.5
SWA-RoPE-LH
35.8
42.6
65.8
33.4
52.3
37.0
25.4
24.8
39.9
27.6
44.8
25.9
37.9
SWA-NoPE-LH
36.4
42.0
66.4
33.4
50.2
37.2
24.1
25.8
39.7
21.6
43.6
25.6
37.2
GLA-RoPE-LH
37.0
42.7
66.4
33.8
52.3
38.8
25.1
24.2
39.0
26.6
44.7
26.3
38.1
Appendix
Table 3: Short-context performance of layer-wise hybrid models after short-context pretraining.
TQA
LMD
PIQA
Hella
Wino
ARC-e
ARC-c
GPQA
SIQA
OBQA
SG
MMLU
Avg.
376M Short
RoPE
36.8
40.2
65.9
33.8
52.2
39.7
26.4
26.3
39.2
27.4
42.1
24.6
37.9
RoPE-NoPE-HH
37.1
40.1
66.0
33.3
50.4
40.0
24.8
27.8
39.9
27.8
43.3
25.8
38.0
SWA-RoPE-HH
35.1
41.3
65.4
33.6
51.3
39.7
23.7
22.7
38.8
28.0
45.0
26.3
37.6
SWA-NoPE-HH
35.8
41.7
65.5
33.8
51.7
39.0
26.4
27.8
38.7
23.4
42.2
25.2
37.6
GLA-RoPE-HH
36.0
46.0
66.0
35.0
52.4
40.2
24.8
25.8
39.1
23.6
43.5
25.6
38.2
Appendix
Table 4: Short-context performance of head-wise hybrid models after short-context pretraining.
TQA
LMD
PIQA
Hella
Wino
ARC-e
ARC-c
GPQA
SIQA
OBQA
SG
MMLU
Avg.
376M Long
RoPE
35.2
35.8
64.3
31.2
51.2
36.7
25.8
24.8
38.5
28.0
43.6
24.2
36.6
RoPE-NoPE-LH
36.0
38.7
64.2
31.7
52.5
37.6
28.8
25.8
39.2
28.2
42.9
24.7
37.5
SWA-RoPE-LH
35.4
39.3
64.2
31.0
52.4
39.2
25.8
23.7
38.3
28.0
45.3
25.2
37.3
+ DRoPE
35.5
39.1
65.0
30.8
51.6
39.0
23.4
24.8
38.1
27.6
44.7
25.7
37.1
SWA-NoPE-LH
35.8
39.6
64.3
31.4
49.0
36.7
25.4
25.3
37.6
26.4
44.4
25.0
36.7
Appendix
Table 5: Short-context performance of layer-wise hybrid models after long-context continual pretraining.
TQA
LMD
PIQA
Hella
Wino
ARC-e
ARC-c
GPQA
SIQA
OBQA
SG
MMLU
Avg.
376M Long
RoPE
35.2
35.8
64.3
31.2
51.2
36.7
25.8
24.8
38.5
28.0
43.6
24.2
36.6
RoPE-NoPE-HH
37.9
36.4
63.6
31.3
50.2
38.5
27.1
26.3
37.8
28.0
44.1
26.1
37.3
SWA-RoPE-HH
35.2
37.6
63.3
31.4
51.0
37.0
22.4
24.8
37.9
28.0
45.1
25.1
36.6
+ DRoPE
35.1
37.1
63.3
31.3
52.0
38.1
23.4
24.8
37.2
27.8
45.1
25.2
36.7
SWA-NoPE-HH
35.5
36.9
63.8
31.5
51.7
37.2
26.1
25.8
37.9
27.8
43.8
23.5
36.8
Appendix
Table 6: Short-context performance of head-wise hybrid models after long-context continual pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Short
RoPE
25.8
0.0
0.0
0.0
0.0
21.7
17.8
8.0
0.4
0.0
0.0
0.0
5.2
6.8
18.3
0.1
6.1
+ NTK
26.5
9.9
0.6
0.1
0.0
21.8
17.8
7.8
8.9
2.2
0.6
0.0
7.4
8.4
18.5
2.8
8.0
+ NTK + Log
26.5
14.0
1.9
0.9
0.5
21.8
17.8
8.0
7.4
7.9
2.2
0.6
8.8
9.4
18.5
4.4
9.1
RoPE-NoPE-LH
29.9
0.1
0.0
0.0
0.0
32.9
19.5
15.7
0.6
0.2
0.0
0.0
6.0
9.8
24.5
0.1
8.2
Appendix
Table 7: Long-context performance of 376M and 776M layer-wise hybrid models after short-context pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
1B Short
RoPE
37.3
1.5
0.0
0.0
0.0
47.4
41.7
28.8
1.7
0.3
0.2
0.1
7.8
17.2
38.8
0.5
13.3
+ NTK
39.3
29.7
6.9
0.5
0.7
47.5
41.6
28.8
21.5
9.0
7.0
2.0
15.4
22.5
39.3
9.7
19.5
+ NTK + Log
39.3
33.4
18.7
8.6
5.8
47.4
41.7
28.8
24.7
20.0
13.4
9.5
21.2
26.5
39.3
16.8
24.3
RoPE-NoPE-LH
46.9
0.1
0.2
0.0
0.0
65.2
51.9
48.1
2.5
0.1
0.0
0.0
9.4
24.0
53.0
0.4
17.9
Appendix
Table 8: Long-context performance of 1B and 3B layer-wise hybrid models after short-context pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Short
RoPE
25.8
0.0
0.0
0.0
0.0
21.7
17.8
8.0
0.4
0.0
0.0
0.0
5.2
6.8
18.3
0.1
6.1
+ NTK
26.5
9.9
0.6
0.1
0.0
21.8
17.8
7.8
8.9
2.2
0.6
0.0
7.4
8.4
18.5
2.8
8.0
+ NTK + Log
26.5
14.0
1.9
0.9
0.5
21.8
17.8
8.0
7.4
7.9
2.2
0.6
8.8
9.4
18.5
4.4
9.1
RoPE-NoPE-HH
33.1
0.3
0.0
0.0
0.0
27.1
28.4
27.0
0.3
0.0
0.0
0.0
6.7
11.8
28.9
0.1
9.7
Appendix
Table 9: Long-context performance of 376M and 776M head-wise hybrid models after short-context pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
1B Short
RoPE
37.3
1.5
0.0
0.0
0.0
47.4
41.7
28.8
1.7
0.3
0.2
0.1
7.8
17.2
38.8
0.5
13.3
+ NTK
39.3
29.7
6.9
0.5
0.7
47.5
41.6
28.8
21.5
9.0
7.0
2.0
15.4
22.5
39.3
9.7
19.5
+ NTK + Log
39.3
33.4
18.7
8.6
5.8
47.4
41.7
28.8
24.7
20.0
13.4
9.5
21.2
26.5
39.3
16.8
24.3
RoPE-NoPE-HH
44.4
1.0
0.1
0.0
0.0
54.9
39.7
35.9
11.4
0.0
0.0
0.0
9.1
20.3
43.7
1.6
15.6
Appendix
Table 10: Long-context performance of 1B and 3B head-wise hybrid models after short-context pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Long
RoPE
31.8
26.0
24.9
10.6
2.4
30.1
25.0
21.7
19.1
10.8
5.0
2.6
19.2
16.3
27.2
12.7
17.5
RoPE-NoPE-LH
35.0
28.0
27.1
14.0
3.7
39.7
30.1
26.2
27.5
30.5
28.2
19.2
21.6
28.8
32.8
22.3
25.8
SWA-RoPE-LH
30.4
28.6
25.8
12.8
6.1
31.1
27.9
24.6
23.3
25.6
15.5
12.6
20.7
22.9
28.5
18.8
22.0
+ DRoPE
28.2
23.6
23.9
15.0
11.0
30.4
25.6
21.8
20.9
20.1
18.2
16.2
20.3
21.9
26.5
18.6
21.2
Appendix
Table 11: Long-context performance of layer-wise hybrid models after long-context continual pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Long
RoPE
31.8
26.0
24.9
10.6
2.4
30.1
25.0
21.7
19.1
10.8
5.0
2.6
19.2
16.3
27.2
12.7
17.5
RoPE-NoPE-HH
39.0
33.3
26.3
9.8
7.3
35.9
25.7
27.1
25.3
24.2
14.1
5.9
23.1
22.6
31.9
18.3
22.8
SWA-RoPE-HH
24.3
19.3
16.4
9.7
6.3
39.9
33.8
27.9
27.0
25.0
17.6
13.5
15.2
26.4
31.5
16.8
21.7
+ DRoPE
22.8
20.8
16.6
16.4
14.3
42.6
30.9
29.1
29.5
27.2
26.2
23.8
18.2
29.9
31.3
21.9
25.0
Appendix
Table 12: Long-context performance of head-wise hybrid models after long-context continual pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Long
SWA-NoPE-LH
35.4
28.1
26.8
22.2
17.1
35.3
30.6
29.2
28.0
24.4
20.7
17.3
25.9
26.5
32.6
23.1
26.3
+ wsz=256
35.1
26.4
24.8
21.7
17.7
35.7
31.0
28.7
28.8
25.5
23.5
19.2
25.2
27.5
32.6
23.5
26.5
+ wsz=512
37.5
26.5
27.4
24.4
20.9
39.1
33.4
29.1
30.4
28.0
23.9
20.1
27.4
29.1
34.8
25.2
28.4
+ wsz=1k
36.8
26.0
26.5
23.3
22.1
37.6
33.7
30.0
32.1
30.7
29.3
21.7
27.0
30.7
34.5
26.5
29.2
Appendix
Table 13: Long-context performance of SWA-NoPE-LH under long-context continual pretraining with larger window sizes (wsz). Larger window sizes exhibit better long-context performance, implying the Short-Window Weariness.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Short
SWA-NoPE-LH
29.1
23.5
19.3
11.0
9.2
31.0
27.2
23.9
20.7
17.1
12.9
12.6
18.4
20.8
27.8
15.8
19.8
+ wsz=256
33.2
30.2
20.8
11.3
8.5
42.3
35.2
29.6
24.9
23.1
17.2
15.3
20.8
26.8
35.1
18.9
24.3
+ wsz=512
30.6
26.1
20.6
11.6
8.3
40.5
35.8
32.4
28.3
24.5
17.8
11.2
19.5
27.2
34.8
18.6
24.0
+ wsz=1k
32.9
25.9
22.8
12.4
4.8
26.6
33.8
30.7
28.8
25.7
23.0
21.4
19.8
27.1
31.0
20.6
24.1
Appendix
Table 14: Long-context performance of SWA-NoPE-LH under short-context pretraining with larger window sizes (wsz). Shorter window sizes exhibit better length extrapolation, implying the Long-Window Laziness.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Long
SWA-RoPE-LH
30.4
28.6
25.8
12.8
6.1
31.1
27.9
24.6
23.3
25.6
15.5
12.6
20.7
22.9
28.5
18.8
22.0
+ wsz=2k
30.9
30.9
28.0
16.9
7.6
26.1
25.9
24.7
22.3
21.8
17.8
14.1
22.9
21.8
26.9
19.9
22.3
+ LongCE
32.1
29.3
29.2
28.7
13.5
37.7
31.7
29.0
31.8
32.5
18.8
15.3
26.5
28.1
32.6
24.9
27.5
+ Both
33.7
34.2
32.3
26.8
12.8
38.7
31.5
28.3
31.2
35.2
21.0
18.4
28.0
29.2
33.1
26.5
28.7
Appendix
Table 15: Long-context performance of SWA layer-wise hybrid models after long-context continual pretraining with extended window sizes (wsz) and LongCE enhancement, both relieving the short-context learning trap effectively.
NIAH-SK1
NIAH-SK2
NIAH-SK3
4k
8k
16k
32k
64k
4k
8k
16k
32k
64k
4k
8k
16k
32k
64k
376M Short
GLA-NoPE-LH
100.0
0.0
0.0
0.0
0.0
100.0
88.0
0.0
0.0
0.0
39.0
35.0
0.0
0.0
0.0
+ Log
100.0
0.0
0.0
0.0
0.0
100.0
94.0
3.0
0.0
0.0
39.0
51.0
0.0
0.0
0.0
+ SWLA
100.0
70.0
79.0
75.0
46.0
100.0
94.0
57.0
9.0
0.0
39.0
34.0
17.0
16.0
0.0
+ SWLA & Log
100.0
70.0
81.0
84.0
84.0
100.0
99.0
97.0
92.0
68.0
39.0
34.0
52.0
70.0
37.0
Appendix
Table 16: Extrapolation comparison of the GLA/GDN-NoPE hybrid model on the NIAH-SK1, SK2, and SK3 tasks. Applying a sliding-window linear attention and a log scale for NoPE attention achieves the best length extrapolation.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Short
RoPE
25.8
0.0
0.0
0.0
0.0
21.7
17.8
8.0
0.4
0.0
0.0
0.0
5.2
6.8
18.3
0.1
6.1
RoPE-NoPE-3:1
33.1
0.3
0.0
0.0
0.0
27.1
28.4
27.0
0.3
0.0
0.0
0.0
6.7
11.8
28.9
0.1
9.7
+ EME
33.1
21.9
15.6
15.8
9.3
27.0
28.4
27.0
27.0
25.3
26.0
22.9
19.1
26.2
28.9
20.5
23.3
RoPE-NoPE-1:1
27.6
0.2
0.0
0.0
0.0
29.8
26.7
21.9
0.1
0.5
0.1
0.1
5.6
11.3
26.5
0.1
8.9
Appendix
Table 17: Long-context performance of head-wise hybrids with different hybrid ratios after short-context pre-training.
SD
MD
Sum
ICL
Syn
Code
Avg.
NQ
Qsp
MF
HG
WQ
Msq
GR
QS
MN
TR
TQ
SS
PC
PR
LCC
Re-P
376M Short
RoPE
0.6
5.0
4.1
1.0
2.2
0.5
2.8
4.4
11.3
2.0
3.2
5.2
1.0
4.5
22.0
18.7
5.5
+ NTK
0.6
5.8
5.6
1.0
3.0
0.8
3.9
8.9
9.6
2.5
3.3
8.8
0.2
2.8
27.2
30.8
7.2
+ NTK + Log
1.4
6.1
6.8
1.6
3.9
1.3
4.3
9.3
9.8
1.5
7.3
11.9
0.5
3.6
26.9
31.3
8.0
RoPE-NoPE-LH
0.4
6.8
4.9
1.3
2.7
0.8
3.8
4.7
12.5
2.5
4.5
2.9
1.6
1.1
21.9
19.0
5.7
Appendix
Table 18: LongBench performance of 376M and 776M layer-wise hybrid models after short-context pretraining.
SD
MD
Sum
ICL
Syn
Code
Avg.
NQ
Qsp
MF
HG
WQ
Msq
GR
QS
MN
TR
TQ
SS
PC
PR
LCC
Re-P
1B Short
RoPE
0.4
9.5
7.1
1.3
5.2
0.4
4.2
5.8
10.7
17.5
9.2
12.7
0.3
4.2
18.5
26.9
8.4
+ NTK
0.9
10.3
9.4
4.4
8.7
1.9
6.7
7.1
8.3
25.0
16.4
24.7
2.7
3.1
24.4
38.4
12.0
+ NTK + Log
2.2
10.6
12.1
5.5
9.3
3.8
6.6
12.3
8.4
29.5
32.0
31.2
3.0
4.0
23.9
37.4
14.5
RoPE-NoPE-LH
0.9
10.9
8.9
2.3
6.4
1.0
6.6
4.8
14.5
14.0
11.4
7.8
2.0
1.1
23.9
28.3
9.0
Appendix
Table 19: LongBench performance of 1B and 3B layer-wise hybrid models after short-context pretraining.
SD
MD
Sum
ICL
Syn
Code
Avg.
NQ
Qsp
MF
HG
WQ
Msq
GR
QS
MN
TR
TQ
SS
PC
PR
LCC
Re-P
376M Short
RoPE
0.6
5.0
4.1
1.0
2.2
0.5
2.8
4.4
11.3
2.0
3.2
5.2
1.0
4.5
22.0
18.7
5.5
+ NTK
0.6
5.8
5.6
1.0
3.0
0.8
3.9
8.9
9.6
2.5
3.3
8.8
0.2
2.8
27.2
30.8
7.2
+ NTK + Log
1.4
6.1
6.8
1.6
3.9
1.3
4.3
9.3
9.8
1.5
7.3
11.9
0.5
3.6
26.9
31.3
8.0
RoPE-NoPE-HH
1.2
6.5
5.1
1.0
3.6
0.7
4.3
5.0
15.0
4.5
5.3
5.0
0.1
2.4
18.7
17.9
6.0
Appendix
Table 20: LongBench performance of 376M and 776M head-wise hybrid models after short-context pretraining.
SD
MD
Sum
ICL
Syn
Code
Avg.
NQ
Qsp
MF
HG
WQ
Msq
GR
QS
MN
TR
TQ
SS
PC
PR
LCC
Re-P
1B Short
RoPE
0.4
9.5
7.1
1.3
5.2
0.4
4.2
5.8
10.7
17.5
9.2
12.7
0.3
4.2
18.5
26.9
8.4
+ NTK
0.9
10.3
9.4
4.4
8.7
1.9
6.7
7.1
8.3
25.0
16.4
24.7
2.7
3.1
24.4
38.4
12.0
+ NTK + Log
2.2
10.6
12.1
5.5
9.3
3.8
6.6
12.3
8.4
29.5
32.0
31.2
3.0
4.0
23.9
37.4
14.5
RoPE-NoPE-HH
0.4
12.3
8.8
1.8
7.0
0.7
5.0
3.5
15.4
16.5
11.1
17.0
1.8
4.5
21.4
31.2
9.9
Appendix
Table 21: LongBench performance of 1B and 3B head-wise hybrid models after short-context pretraining.
SD
MD
Sum
ICL
Syn
Code
Avg.
NQ
Qsp
MF
HG
WQ
Msq
GR
QS
MN
TR
TQ
SS
PC
PR
LCC
Re-P
376M Long
RoPE
1.6
9.9
10.3
3.5
6.2
2.6
20.3
10.9
16.8
18.5
17.0
17.2
2.7
3.7
29.5
32.6
12.7
RoPE-NoPE-LH
1.5
12.7
12.4
3.8
7.4
3.0
19.5
10.7
18.9
38.0
19.8
9.6
3.2
3.7
27.8
31.8
14.0
SWA-RoPE-LH
1.6
9.4
11.2
4.6
7.6
3.1
14.5
8.1
15.3
40.0
17.3
16.7
3.4
3.2
34.9
36.8
14.2
+ DRoPE
1.4
9.0
11.5
3.8
7.1
2.9
12.7
9.7
14.2
43.0
18.5
12.5
3.1
3.9
32.1
34.9
13.8
Appendix
Table 22: LongBench performance of layer-wise hybrid models after long-context continual pretraining.
SD
MD
Sum
ICL
Syn
Code
Avg.
NQ
Qsp
MF
HG
WQ
Msq
GR
QS
MN
TR
TQ
SS
PC
PR
LCC
Re-P
376M Long
RoPE
1.6
9.9
10.3
3.5
6.2
2.6
20.3
10.9
16.8
18.5
17.0
17.2
2.7
3.7
29.5
32.6
12.7
RoPE-NoPE-HH
2.7
9.8
12.3
4.0
7.4
2.6
22.6
9.1
21.7
31.0
18.2
14.5
2.4
3.8
29.2
32.4
14.0
SWA-RoPE-HH
1.9
11.3
12.9
3.8
6.3
3.2
18.2
13.8
16.5
23.0
17.0
15.1
3.0
3.4
29.6
35.0
13.4
+ DRoPE
1.7
9.5
12.7
4.1
6.3
3.1
18.6
12.8
17.0
26.0
15.8
13.3
3.1
2.9
34.3
35.9
13.6
Appendix
Table 23: LongBench performance of head-wise hybrid models after long-context continual pretraining.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Short
RoPE
25.8
0.0
0.0
0.0
0.0
21.7
17.8
8.0
0.4
0.0
0.0
0.0
5.2
6.8
18.3
0.1
6.1
RoPE-NoPE-LH
28.3
0.1
0.0
0.0
0.0
41.2
27.1
24.8
0.4
0.1
0.0
0.0
5.7
13.4
30.3
0.1
10.2
SWA-NoPE-LH
23.2
16.0
14.0
13.2
10.5
31.8
26.1
21.8
15.8
14.7
12.0
10.6
15.4
19.0
25.7
13.4
17.5
GLA-NoPE-LH
21.0
3.8
1.5
2.3
2.2
41.5
36.4
33.5
20.9
16.3
16.4
17.0
6.2
26.0
33.1
10.1
17.7
Appendix
Table 24: Long-context performance of layer-wise hybrid models under the YOCO-like layout.
RULER
BABILong
Average
4k
8k
16k
32k
64k
0k
2k
4k
8k
16k
32k
64k
RU.
BA
≤ 4k
¿4k
All
376M Short
SWA-NoPE-LH
29.1
23.5
19.3
11.0
9.2
31.0
27.2
23.9
20.7
17.1
12.9
12.6
18.4
20.8
27.8
15.8
19.8
+ base=5000
27.9
19.6
12.7
8.0
7.0
40.7
34.5
32.1
27.5
24.0
15.9
14.2
15.0
27.0
33.8
16.1
22.0
+ base=2500
28.5
16.2
11.4
9.1
7.7
45.7
33.8
28.0
24.0
17.2
11.3
14.6
14.6
24.9
34.0
13.9
20.6
+ base=1000
28.7
21.9
15.5
10.9
7.7
32.6
32.0
28.8
25.4
21.2
13.1
11.4
16.9
23.5
30.5
15.9
20.8
Appendix
Table 25: Long-context performance of SWA-NoPE-LH with smaller rotary bases, from 10000 to 100, under direct or log-scaled NoPE extrapolation after short-context pretraining, and context extension after long-context pretraining.
Modern language models increasingly adopt hybrid architectures that combine full attention with efficient attention modules, such as sliding-window attention (SWA) and recurrent sequence mixers. However, how these efficient modules shape model capabilities remains poorly understood. To address this gap, we conduct a systematic analysis across hybrid architectures from three perspectives: scaling behavior, mechanism analysis, and architecture design. First, from a scaling perspective, we find that efficient-attention design primarily affects how fast long-context capability emerges, while different hybrids eventually converge to comparable long-context performance under sufficient training. Second, mechanistically, we show that long-range retrieval is mainly carried by full attention, whereas efficient attention shapes its optimization trajectory. This explains a counter-intuitive phenomenon we call Large-Window Laziness: larger SWA windows can delay the formation of retrieval heads in full-attention layers. Third, guided by this mechanism, we show that applying NoPE to only the full-attention layers of a small-window SWA hybrid substantially improves long-context performance with negligible impact on short-context performance.
The quadratic complexity of attention poses a critical bottleneck for long-context processing, spurring interest in hybrid attention designs. Most open-source hybrid models adopt a layer-wise strategy. Yet, prior work has noted the inherent difficulty of integrating Linear Attention (LA) with Full Attention (FA), suggesting that the design space of attention hybridization remains underexplored. To probe this space, we conduct interpretability analysis and observe that layers exhibit block-wise functional similarity, while individual heads within the same layer display distinct functional specialization despite sharing input features. This head-level heterogeneity suggests that the head dimension provides a natural and principled granularity for fusing heterogeneous attention signals. Building on this insight, we introduce HydraHead, a novel architecture that hybridizes FA and LA along the head axis. HydraHead features two key innovations: (1) an interpretability-driven selection strategy that identifies retrieval-critical heads and preserves FA only for them, and (2) a scale-normalized fusion module that reconciles the distributional gap between FA and LA head outputs. By leveraging a three-stage transfer pipeline with parameter reuse and distillation, we achieve high-performance hybrid models with minimal training overhead. Under a unified training setup, HydraHead outperforms other hybrid designs in long-context tasks while maintaining strong general reasoning. With interpretability-driven head selection, it matches a 3:1 layer-wise hybrid's long-context performance at a 7:1 LA-to-FA ratio. Crucially, trained on only 15B tokens, HydraHead achieves over 69% improvement over the baseline at 512K context length, approaching Qwen3.5, a leading model of comparable size with a native context length of 256K. This highlights the significant scaling potential of head-level hybridization.
Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued pretraining on long-context data can help but is expensive due to the quadratic scaling of Attention. We observe that most tokens do not require (Global) Attention over the entire sequence and can rely on local context. Based on this, we propose L2A (Learning To Attend), a layer that enables conditional (token-wise) long-range memory access by deciding when to invoke global attention. We evaluate L2A on Qwen 2.5 and Qwen 3 models, extending their effective context length from 32K to 128K tokens. L2A matches the performance of standard long-context training to within 3% while skipping Global Attention for ∼80% of tokens, outperforming prior baselines. We also design custom Triton kernels to efficiently implement this token-wise conditional Attention on GPUs, achieving up to ∼2× improvements in training throughput and time-to-first-token over FlashAttention. Moreover, L2A enables post-training pruning of highly sparse Global Attention layers, reducing KV cache memory by up to 50% with negligible performance loss. Our code is released under Apache 2.0 at https://github.com/awslabs/hybrid-model-factory/tree/main/examples/research/L2A.