Multilinguality in Hybrid Attention LLMs
Organizations: University of California, Los Angeles · Fudan University
Abstract
In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of LLMs. Beyond the impact on long sequences in poorly tokenized languages, our study is motivated by the possibility that the inductive biases of the recurrent state alter linguistic processing. Our interpretability analysis confirms this, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers. Across diverse models, we notably observe a pronounced spike in cross-lingual alignment around the first full-attention layer. These findings lead us to question the conventional ordering of attention layers. In distillation experiments on multilingual data, all alternative layer orderings outperform the standard throughout training, learning up to 2.5X faster. These stark, replicable results prompt our theory that multilingual models would benefit from starting with a full-attention layer rather than recurrent layers.
Figures & tables
| Hybrid-Attn Model | Recurrent Variant | Ratio | Full Attn Variant | Comparable Non-Hybrid | Relationship |
| Qwen3.5-35B-A3B | Gated DeltaNet | 3:1 | Gated GQA | Qwen3-30B-A3B | Unknown, same architecture outside attention |
| OLMo-Hybrid | Gated DeltaNet w/ neg. eigenvalues | 3:1 | GQA | OLMo-3-7B a | Exact same data recipe, trainings from scratch controlled for comparability. |
| Ring-mini-linear-2.0 | Lightning Attn 2 | 4:1 | MLA | Ring-mini-2.0 | Same base model, with hybrid conversion via distillation happening before comparable post-trainings |
| Granite-4.0-h-micro | Mamba-2 | 9:1 b | GQA | Granite-4.0-Micro | Same data recipe, each trained from scratch, unknown how controlled |
| Ling-2.6-flash | Lightning Attn 2 | 7:1 | MLA | — | — |
| Tiny Aya Global | SWA | 3:1 | GQA | — | — |
| English | Non-Eng (10 langs) | |||||
|---|---|---|---|---|---|---|
| Model | 8K | 32K | 64K | 8K | 32K | 64K |
| Qwen3-30B-A3B | 99.2 | 98.0 | 97.6 | 98.2 | 95.3 | 92.0 |
| Qwen3.5-35B-A3B | 100.0 | 100.0 | 99.6 | 96.9 | 95.6 | 93.8 |
| Granite-4.0-Micro | 80.4 | 65.6 | 47.2 | 71.2 | 47.9 | 38.1 |
| Granite-4.0-H-Micro | 75.2 | 68.0 | 61.6 | 52.1 | 39.5 | 34.9 |
| Teacher | Placement | relative | Gap |
| Qwen3-4B-Base | Standard Periodic | 1.000 / 1.000 | 0.071 / 0.068 |
| ”” | Reverse Periodic | 0.677 / 0.697 | 0.045 / 0.045 |
| ”” | Sandwich | 0.665 | 0.050 |
| ”” | Interleaved Sandwich | 0.689 | 0.049 |
| ”” | Endpoint Spread | 0.722 | 0.049 |
| Granite-4.1-3B-Base | Standard Periodic | 1.000 | 0.042 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Model Name | Citation | # Layers | MoE ? | Params | Hidden Size | Position |
| Qwen3.5-35B-A3B | Qwen Team (2026) | 40 | Yes | 35B | 2048 | RoPE |
| Qwen3-30B-A3B | Yang et al. (2025a) | 48 | ”” | 30B | ”” | ”” |
| OLMo-Hybrid | Merrill et al. (2026) | 32 | No | 7B | 3840 | RoPE |
| OLMo-3-7B | Olmo: et al. (2026) | ”” | ”” | ”” | 4096 | RoPE & YaRN |
| Ring-mini-linear-2.0 | Ling Team et al. (2025a) | 20 | Yes | 16B | 2048 | RoPE |
| Ring-mini-2.0 | Ling Team et al. (2025b) | ”” | ”” | ”” | ”” | ”” |
| Display Name | Full Name | Citation |
|---|---|---|
| GQA | Grouped Query Attention | Ainslie et al. (2023) |
| Gated GQA | — | Qiu et al. (2025) |
| MLA | Multi-Head Latent Attention | DeepSeek-AI et al. (2024) |
| SWA | Sliding Window Attention | Beltagy et al. (2020) |
| Gated DeltaNet | — | Yang et al. (2025d) |
| Gated DeltaNet w/ Negative Eigenvalues | — | Grazzi et al. (2025) |
| Tier | Languages | Share Each | Tokens Each |
|---|---|---|---|
| English | eng_Latn | 30.0% | 300M |
| High (8) | rus_Cyrl, hin_Deva, cmn_Hani, arb_Arab | 4.08% | 40.8M |
| ind_Latn, vie_Latn, tur_Latn, tam_Taml | |||
| Mid (14) | ell_Grek, hye_Armn, yue_Hani, mya_Mymr, heb_Hebr | 2.33% | 23.3M |
| amh_Ethi, hau_Latn, zsm_Latn, fil_Latn, tel_Telu | |||
| mal_Mlym, azj_Latn, kaz_Cyrl, khm_Khmr |
| Family | Language | Code | Script Type | Tier | FW2 Doc # |
|---|---|---|---|---|---|
| Indo-European | English | eng_Latn | Alphabet | Anchor | — |
| Russian | rus_Cyrl | Alphabet | High | 699.1M | |
| Hindi | hin_Deva | Abugida | High | 22.1M | |
| Greek | ell_Grek | Alphabet | Mid | 47.4M | |
| Armenian | hye_Armn | Alphabet | Mid | 1.8M | |
| Irish | gle_Latn | Alphabet | Low | 0.65M |
| Hyperparameter | Qwen3-4B-Base | Granite-4.1-3B-Base |
|---|---|---|
| Yang et al. (2025a) | IBM Research (2026) | |
| # Layers (full / recurrent) | 36 (9 / 27) | 40 (10 / 30) |
| Hidden size | 2560 | 2560 |
| Heads / | 32 / 8 | 40 / 8 |
| Head dim | 128 | 64 |
| Recurrent state per layer ( ) | 524K | 164K |
| Qwen3-4B | Granite-4.1-3B | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| data seed 0 | data seed 1 | data seed 1 | |||||||||
| Language | Tier | Metric | SP | RP | SW | IS | ES | SP | RP | SP | RP |
| English | E | 0.160 | 0.140 | 0.148 | 0.147 | 0.144 | 0.160 | 0.142 | 0.212 | 0.203 | |
| 0.119 | 0.106 | 0.116 | 0.112 | 0.108 | 0.120 | 0.108 | 0.157 | 0.151 | |||
| Russian | H | 0.113 | 0.092 | 0.100 | 0.099 | 0.096 | 0.111 | 0.093 | 0.099 | 0.097 | |
| 0.088 | 0.073 | 0.082 | 0.079 | 0.076 | 0.089 | 0.074 | 0.054 | 0.053 | |||