In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of LLMs. Beyond the impact on long sequences in poorly tokenized languages, our study is motivated by the possibility that the inductive biases of the recurrent state alter linguistic processing. Our interpretability analysis confirms this, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers. Across diverse models, we notably observe a pronounced spike in cross-lingual alignment around the first full-attention layer. These findings lead us to question the conventional ordering of attention layers. In distillation experiments on multilingual data, all alternative layer orderings outperform the standard throughout training, learning up to 2.5X faster. These stark, replicable results prompt our theory that multilingual models would benefit from starting with a full-attention layer rather than recurrent layers.
Figures & tables
Hybrid-Attn Model
Recurrent Variant
Ratio
Full Attn Variant
Comparable Non-Hybrid
Relationship
Qwen3.5-35B-A3B
Gated DeltaNet
3:1
Gated GQA
Qwen3-30B-A3B
Unknown, same architecture outside attention
OLMo-Hybrid
Gated DeltaNet w/ neg. eigenvalues
3:1
GQA
OLMo-3-7B a
Exact same data recipe, trainings from scratch controlled for comparability.
Ring-mini-linear-2.0
Lightning Attn 2
4:1
MLA
Ring-mini-2.0
Same base model, with hybrid conversion via distillation happening before comparable post-trainings
Granite-4.0-h-micro
Mamba-2
9:1 b
GQA
Granite-4.0-Micro
Same data recipe, each trained from scratch, unknown how controlled
Ling-2.6-flash
Lightning Attn 2
7:1
MLA
—
—
Tiny Aya Global
SWA
3:1
GQA
—
—
Table 1: Hybrid-attention models studied in this work and their closest non-hybrid counterparts. In the interest of space, we provide further design details and citations for all in Appendix A . The much larger Ling-2.6-flash and SWA-hybrid Tiny Aya are also studied partially.
English
Non-Eng (10 langs)
Model
8K
32K
64K
8K
32K
64K
Qwen3-30B-A3B
99.2
98.0
97.6
98.2
95.3
92.0
Qwen3.5-35B-A3B
100.0
100.0
99.6
96.9
95.6
93.8
Granite-4.0-Micro
80.4
65.6
47.2
71.2
47.9
38.1
Granite-4.0-H-Micro
75.2
68.0
61.6
52.1
39.5
34.9
Table 2: OneRULER accuracy ↑ (%) on needle-in-a-haystack (NIAH) tasks at different context sizes. Note that Qwen3.5 is 9 months newer, bigger, and better than Qwen3.
Figure 1: Formalization of the decoder layer structure applicable for all decoder layers across models.
Figure 2: SoftCKA scores for Granite models, with hidden states taken at the beginning of each decoder layer ( attn_in ). Red dotted lines mark where full attention layers are in Granite H.
Figure 3: Visualization of the jagged curves in the hybrid models (right), which seem to reflect the heterogeneous attention layout, compared with the smoother curves in non-hybrid counterparts(left).
Figure 4: SoftCKA scores of the attention block delta . Curiously, the major “event” at the first full attention layer is opposite to other models.
Figure 5: Full attention placement schemas at a fixed 25% budget to maintain the 3:1 ratio between recurrent attention and full attention (9/36 layers for Qwen3-4B, 10/40 for Granite-4.1-3B). Tall bars = full attention, short grey ticks = Gated DeltaNet; layer 0 is closest to the embeddings. Students differ only in placement, not in how many full-attention layers they keep.
Figure 6: Held-out LLM gap from the teacher (nats/token, mean over 26 languages, validation every 10M tokens). (a) Qwen3-4B, all five placements, data seed 0. (b) Granite-4.1-3B, standard periodic and reverse periodic , data seed 1. The red dashed line marks standard periodic ’s value at the last checkpoint (0.073 nats in (a) and 0.033 in (b) at 990M tokens). In (a), all alternatives reach it within ∼40% of training, with reverse periodic at 34.5%. In (b), reverse periodic reaches it at 42.5%.
Teacher
Placement
relative DKLR↓
LLM Gap δˉ↓
Qwen3-4B-Base
Standard Periodic
1.000 / 1.000
0.071 / 0.068
””
Reverse Periodic
0.677 / 0.697
0.045 / 0.045
””
Sandwich
0.665
0.050
””
Interleaved Sandwich
0.689
0.049
””
Endpoint Spread
0.722
0.049
Granite-4.1-3B-Base
Standard Periodic
1.000
0.042
Table 3: Test results after 1B tokens, averaged over 26 languages. R : DKL to the teacher relative to standard periodic with the same teacher and training set (equation 5 ); δˉ : mean LLM gap to the teacher in nats (equation 4 ). For Qwen3-4B-Base, standard periodic and reverse periodic are trained twice on independently resampled data, shown as run 1 / run 2.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Model Name
Citation
# Layers
MoE ?
Params
Hidden Size
Position
Qwen3.5-35B-A3B
Qwen Team (2026)
40
Yes
35B
2048
RoPE
Qwen3-30B-A3B
Yang et al. (2025a)
48
””
30B
””
””
OLMo-Hybrid
Merrill et al. (2026)
32
No
7B
3840
RoPE
OLMo-3-7B
Olmo: et al. (2026)
””
””
””
4096
RoPE & YaRN
Ring-mini-linear-2.0
Ling Team et al. (2025a)
20
Yes
16B
2048
RoPE
Ring-mini-2.0
Ling Team et al. (2025b)
””
””
””
””
””
Appendix
Table A.1: Continuing on from Table 1 , we provide additional architecture details. “Position” is the positional embeddings used in the full attention layers .
Display Name
Full Name
Citation
GQA
Grouped Query Attention
Ainslie et al. (2023)
Gated GQA
—
Qiu et al. (2025)
MLA
Multi-Head Latent Attention
DeepSeek-AI et al. (2024)
SWA
Sliding Window Attention
Beltagy et al. (2020)
Gated DeltaNet
—
Yang et al. (2025d)
Gated DeltaNet w/ Negative Eigenvalues
—
Grazzi et al. (2025)
Appendix
Table A.2: Architectural Components
Tier
Languages
Share Each
Tokens Each
English
eng_Latn
30.0%
300M
High (8)
rus_Cyrl, hin_Deva, cmn_Hani, arb_Arab
4.08%
40.8M
ind_Latn, vie_Latn, tur_Latn, tam_Taml
Mid (14)
ell_Grek, hye_Armn, yue_Hani, mya_Mymr, heb_Hebr
2.33%
23.3M
amh_Ethi, hau_Latn, zsm_Latn, fil_Latn, tel_Telu
mal_Mlym, azj_Latn, kaz_Cyrl, khm_Khmr
Appendix
Table C.1: Distillation mixture ( N=109 tokens). Shares are token quotas under each teacher’s tokenizer. English is drawn from FineWeb (sample-100BT) and all other languages from FineWeb-2; codes follow FineWeb-2 (ISO 639-3 and script).
Family
Language
Code
Script Type
Tier
FW2 Doc #
Indo-European
English
eng_Latn
Alphabet
Anchor
—
Russian
rus_Cyrl
Alphabet
High
699.1M
Hindi
hin_Deva
Abugida
High
22.1M
Greek
ell_Grek
Alphabet
Mid
47.4M
Armenian
hye_Armn
Alphabet
Mid
1.8M
Irish
gle_Latn
Alphabet
Low
0.65M
Appendix
Table C.2: Distillation languages, grouped by family. FW2 Docs is the number of FineWeb-2 training documents ( Penedo et al., 2025 ) ; English is drawn from FineWeb ( Penedo et al., 2024 ) . Tier sets each language’s share of the 109 -token budget: English 30%, High 4.08%, Mid 2.33%, Low 1.17% (a 3.5 : 2 : 1 ratio), i.e., 300M, 40.8M, 23.3M, and 11.7M tokens under each teacher’s tokenizer.
Hyperparameter
Qwen3-4B-Base
Granite-4.1-3B-Base
Yang et al. (2025a)
IBM Research (2026)
# Layers (full / recurrent)
36 (9 / 27)
40 (10 / 30)
Hidden size d
2560
2560
Heads H / Hkv
32 / 8
40 / 8
Head dim dh
128
64
Recurrent state per layer ( Hdh2 )
524K
164K
Appendix
Table C.3: Distillation hyperparameters and experimental details. Symbols follow Algorithm 1 and Eq. equation 3 .
Figure C.1: Distillation training curves for Qwen3-4B with data seed 0, one run per placement. Faint lines show the per-step DKL on the training batch and bold lines its exponential moving average ( α=0.1 ). The y-axis is logarithmic and cut off at 0.5 nats; every run starts between 7.0 and 10.5 nats.
Figure C.2: Distillation training curves for the Qwen3-4B and Granite-4.1-3B standard periodic and reverse periodic placement with data seed 1.
Qwen3-4B
Granite-4.1-3B
data seed 0
data seed 1
data seed 1
Language
Tier
Metric
SP
RP
SW
IS
ES
SP
RP
SP
RP
English
E
DKL
0.160
0.140
0.148
0.147
0.144
0.160
0.142
0.212
0.203
LLM
0.119
0.106
0.116
0.112
0.108
0.120
0.108
0.157
0.151
Russian
H
DKL
0.113
0.092
0.100
0.099
0.096
0.111
0.093
0.099
0.097
LLM
0.088
0.073
0.082
0.079
0.076
0.089
0.074
0.054
0.053
Appendix
Table C.4: Per-language test results. For each language, the DKL row gives DKL(pT∥pS) (nats per token) and the LLM row the LLM gap to the teacher, LLM,λS−LLM,λT (nats); lower is better for both, and a negative LLM gap means the student is less perplexed than the teacher. The last two rows give R (equation 5 ), relative to SP of the run with the same data seed, and δˉ (equation 4 ). Teachers are the base models (Qwen3-4B-Base, Granite-4.1-3B-Base). Languages are grouped by family as in Table C.2 . SP: Standard Periodic ; RP: Reverse Periodic ; SW: Sandwich ; IS: Interleaved Sandwich ; ES: Endpoint Spread . Tier: E/H/M/L = English/High/Mid/Low sampling tier (Table C.2 ). † Mon is excluded from all averages.
Figure E.1: SoftCKA of the residual-stream contributions from the attention and MLP blocks in Qwen3.5. The blocks preceding the first full-attention layer show a progressive change before the pronounced event at that layer. Similarity has a massive drop right after this layer.
Figure F.1: SoftCKA for Tiny Aya Global, which is an SWA-hybrid model. Red dotted lines mark full-attention layers.
Figure G.1: Layer-wise cross-lingual MoE routing divergence for Qwen3.5 (hybrid). Lower values indicate more similar routing across languages. Red dotted lines mark the full attention layers.
Figure G.2: Layer-wise cross-lingual MoE routing divergence for Qwen3 (non-hybrid).
Figure G.3: Layer-wise cross-lingual MoE routing divergence for Ring-mini-linear-2.0 (hybrid). Red dotted lines mark the full attention layers.
Figure G.4: Layer-wise cross-lingual MoE routing divergence for Ring-mini-2.0 (non-hybrid).
Hybrid language models that interleave attention with recurrent components are increasingly competitive with pure Transformers, yet standard LoRA practice applies adapters uniformly without considering the distinct functional roles of each component type. We systematically study component-type LoRA placement across two hybrid architectures -- Qwen3.5-0.8B (sequential, GatedDeltaNet + softmax attention) and Falcon-H1-0.5B (parallel, Mamba-2 SSM + attention) -- fine-tuned on three domains and evaluated on five benchmarks. We find that the attention pathway -- despite being the minority component -- consistently outperforms full-model adaptation with 5-10x fewer trainable parameters. Crucially, adapting the recurrent backbone is destructive in sequential hybrids (-14.8 pp on GSM8K) but constructive in parallel ones (+8.6 pp). We further document a transfer asymmetry: parallel hybrids exhibit positive cross-task transfer while sequential hybrids suffer catastrophic forgetting. These results establish that hybrid topology fundamentally determines adaptation response, and that component-aware LoRA placement is a necessary design dimension for hybrid architectures.
VRAIN – Valencian Research Institute for Artificial Intelligence, Universitat Politècnica de València, Valencia, Spain · Department of Economics and Social Sciences, Universitat Politècnica de València, Valencia, Spain · Department of Business Organisation, Universitat Politècnica de València, Valencia, Spain
Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and recurrent layers offer complementary pathways to use past information: attention supports precise memory recall from earlier tokens, while recurrent layers support consolidation of disparate information over long contexts. However, we observe that simply having access to both pathways does not mean that hybrid LMs are effectively using them. We find that they rely substantially more on attention than on the recurrent state. Standard supervised fine-tuning improves overall performance but does not improve how the two memory pathways are coordinated: the model becomes more reliant on information propagated by attention layers, while its use of information propagated by recurrent layers remains limited. To encourage better coordination between the two memory pathways, we add an auxiliary loss that limits attention's access to earlier context while the recurrent state propagates through the full sequence. This objective encourages the model to retain and use information through the recurrent pathway alongside attention. It improves overall performance, with particularly strong gains on tasks involving longer contexts or requiring information aggregation, consistent with the strengths of recurrent layers observed in analysis. Crucially, this imbalance and the benefit of our auxiliary loss generalize: they apply to multiple recurrent-attention LMs in question-answering and agentic tasks, as well as to attention-based LMs that combine different forms of memory. Together, our findings show that simply providing multiple memory pathways does not ensure their effective use, and that targeted supervision is needed to better coordinate them.
Hyunji Lee, Joykirat Singh, Zaid Khan +5
UNC Chapel Hill · The University of Texas at Austin · Mila +1
The quadratic complexity of attention poses a critical bottleneck for long-context processing, spurring interest in hybrid attention designs. Most open-source hybrid models adopt a layer-wise strategy. Yet, prior work has noted the inherent difficulty of integrating Linear Attention (LA) with Full Attention (FA), suggesting that the design space of attention hybridization remains underexplored. To probe this space, we conduct interpretability analysis and observe that layers exhibit block-wise functional similarity, while individual heads within the same layer display distinct functional specialization despite sharing input features. This head-level heterogeneity suggests that the head dimension provides a natural and principled granularity for fusing heterogeneous attention signals. Building on this insight, we introduce HydraHead, a novel architecture that hybridizes FA and LA along the head axis. HydraHead features two key innovations: (1) an interpretability-driven selection strategy that identifies retrieval-critical heads and preserves FA only for them, and (2) a scale-normalized fusion module that reconciles the distributional gap between FA and LA head outputs. By leveraging a three-stage transfer pipeline with parameter reuse and distillation, we achieve high-performance hybrid models with minimal training overhead. Under a unified training setup, HydraHead outperforms other hybrid designs in long-context tasks while maintaining strong general reasoning. With interpretability-driven head selection, it matches a 3:1 layer-wise hybrid's long-context performance at a 7:1 LA-to-FA ratio. Crucially, trained on only 15B tokens, HydraHead achieves over 69% improvement over the baseline at 512K context length, approaching Qwen3.5, a leading model of comparable size with a native context length of 256K. This highlights the significant scaling potential of head-level hybridization.