Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and recurrent layers offer complementary pathways to use past information: attention supports precise memory recall from earlier tokens, while recurrent layers support consolidation of disparate information over long contexts. However, we observe that simply having access to both pathways does not mean that hybrid LMs are effectively using them. We find that they rely substantially more on attention than on the recurrent state. Standard supervised fine-tuning improves overall performance but does not improve how the two memory pathways are coordinated: the model becomes more reliant on information propagated by attention layers, while its use of information propagated by recurrent layers remains limited. To encourage better coordination between the two memory pathways, we add an auxiliary loss that limits attention's access to earlier context while the recurrent state propagates through the full sequence. This objective encourages the model to retain and use information through the recurrent pathway alongside attention. It improves overall performance, with particularly strong gains on tasks involving longer contexts or requiring information aggregation, consistent with the strengths of recurrent layers observed in analysis. Crucially, this imbalance and the benefit of our auxiliary loss generalize: they apply to multiple recurrent-attention LMs in question-answering and agentic tasks, as well as to attention-based LMs that combine different forms of memory. Together, our findings show that simply providing multiple memory pathways does not ensure their effective use, and that targeted supervision is needed to better coordinate them.
Figures & tables
Figure 1: Overview of memory pathways in recurrent-attention hybrid LMs and our findings. ( Middle ) Two memory pathways in recurrent-attention hybrid LMs play different roles: attention supports precise memory recall of earlier tokens, while the recurrent state supports aggregating information across long contexts. ( Right ) With Standard SFT ( LSFT ), the model relies mostly on attention and underuses the recurrent pathways. We add an auxiliary loss ( Lrec ) that masks attention to past context, allowing the model to access it only through recurrent layers. We observe that adding auxiliary loss improves overall performance, with increased use of the recurrent pathway and preserved use of attention.
Figure 2Figure 3
Model
Training Obj.
DROP
MuSR
HQA
Qasper
NQA
MINT
Avg.
Qwen3.5
LSFT
48.6
45.6
43.1
30.1
24.4
21.1
34.6
LSFT+Lrec (Ours)
46.9
44.8
44.0
35.9
33.7
28.9
38.5
Lrec
35.0
36.5
38.0
32.1
20.1
23.6
30.4
Lattn
56.9
48.1
35.7
31.3
15.6
10.3
31.6
LSFT+Lattn
54.0
49.1
45.7
30.3
23.7
17.1
35.6
Nemotron-H
LSFT
42.4
40.9
38.7
28.1
21.9
18.8
31.0
Table 1: QA accuracy across six datasets. First two rows compare standard SFT with our objective; the remaining rows ablate the training objective. Avg. is micro-average over all test examples.
Figure 6: QA accuracy of Qwen3.5-4B across various training objectives (details of objectives in Section 5.2 ). Hatched bars indicate that it adds an auxiliary loss. SFT+REC increases Recurrent-only accuracy while largely preserving Attention-only accuracy and best Full accuracy.
Model
Training Obj.
TextWorld
BabyAI
Quest
Treasure
PutNextLocal
GoToObjMaze
Qwen3.5
LSFT
78.5
44.0
46.0
63.0
LSFT+Lrec (Ours)
80.5
56.0
73.0
83.0
Lrec
75.5
45.0
13.0
37.0
Lattn
76.5
44.0
31.0
50.0
LSFT+Lattn
77.0
49.0
47.0
70.0
Table 2: Performance on TextWorld and BabyAI under different training objectives. For each model, the first two rows compare standard SFT with our objective; the remaining rows ablate the training objective. Best in bold .
Figure 7: Auxiliary supervision improves learning from past trajectories. Left: Cumulative recovery rate on QA as the maximum number of attempts increases, with previous failed attempts provided in context. The auxiliary loss improves recovery, especially on multi-evidence questions. Right: Rate of generating trajectories shorter than all provided successful demonstrations, grouped by median demonstration length relative to the optimal path. The auxiliary loss yields shorter trajectories more often, especially with longer demonstrations.
Figure 8: QA accuracy of attention-based models given a single memory type ( Text , Latent ), both ( SFT , +Aux ), or per-example Oracle upper bound.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9
Figure 11: Relative performance of latent versus textual memory (y-axis) under matched memory budgets (x-axis). Textual memory performs better on verbatim and single-evidence questions, while latent memory is more competitive on non-verbatim and multi-evidence questions.
Role-conditioned
Variant
Acc.
Δ Latent
Δ Text
-
Naive
24.1
-2.4
+11.4
Yes
Outcome
26.1
+0.8
+11.5
Verbatim
28.6
+4.7
+10.4
Reverse
24.5
+1.1
+7.2
Random
23.9
-1.0
+10.5
No
Latent-aux
26.1
+3.4
+8.5
Appendix
Table 3: Accuracy and memory use across auxiliary-training variants. Δ Latent and Δ Text denote the accuracy drop when the latent or textual memory, respectively, is replaced with one constructed from an unrelated context; higher values indicate greater use of that memory. Role-conditioned indicates whether auxiliary supervision uses per-example role assignment.
Transformers face quadratic complexity and memory issues with long sequences, prompting the adoption of linear attention mechanisms using fixed-size hidden states. However, linear models often suffer from limited recall performance, leading to hybrid architectures that combine linear and full attention layers. Despite extensive hybrid architecture research, the choice of linear attention component has not been deeply explored. We systematically evaluate various linear attention models across generations - vector recurrences to advanced gating mechanisms - both standalone and hybridized. To enable this comprehensive analysis, we trained and open-sourced 72 models: 36 at 340M parameters (20B tokens) and 36 at 1.3B parameters (100B tokens), covering six linear attention variants across five hybridization ratios. Benchmarking on standard language modeling and recall tasks reveals that superior standalone linear models do not necessarily excel in hybrids. While language modeling remains stable across linear-to-full attention ratios, recall significantly improves with increased full attention layers, particularly below a 3:1 ratio. Our study highlights selective gating, hierarchical recurrence, and controlled forgetting as critical for effective hybrid models. We recommend architectures such as HGRN-2 or GatedDeltaNet with a linear-to-full ratio between 3:1 and 6:1 to achieve Transformer-level recall efficiently. Our models are open-sourced at https://huggingface.co/collections/m-a-p/hybrid-linear-attention-research-686c488a63d609d2f20e2b1e.
Dustin Wang, Rui-Jie Zhu, Steven Abreu +9
1UC Santa Cruz · 2ByteDance Seed · University of Groningen +3
Softmax attention is the cornerstone of modern large language models, but its memory scales linearly and compute quadratically with sequence length. Linear recurrent models, such as linear attention and state space models, have become widely studied as alternatives to attention due to their linear compute and constant memory. While these sub-quadratic token mixing methods, or mixers, achieve promising efficiency gains and competitive results on a wide range of benchmarks, current linear recurrent models still lag behind on tasks that require long-context retrieval or in-context learning. A growing body of work studies hybrid architectures that attempt to mitigate these trade-offs by statically interleaving or merging attention and recurrent blocks. In this work, we explore a new axis of developing hybrid models: across the token sequence. We propose Oryx, a hybrid model that can, throughout a sequence, flexibly switch between different mixers, for example quadratic attention for rich context utilization and linear recurrences for efficient generation. Oryx ties at least 90% of its parameters across mixers, enabling attention and recurrent modes to operate over shared internal representations. We validate our design with Mamba-2 and Gated DeltaNet variants, up to 1.4B models. Under fixed token budgets and a mixed-training strategy, Oryx achieves comparable or better performance than its single-mixer baselines. At the 1.4B scale, all instances of Oryx outperform their respective baselines by at least 0.7 percentage points on averaged language modeling tasks. On retrieval tasks, Oryx achieves performance comparable to the Transformer baseline even when processing only a tiny fraction (<10%) of the tokens in attention mode. These results suggest that attention and linear recurrent models can share internal representations, and motivate sequence-axis hybridization as a promising direction.
Kevin Y. Li, Asher Trockman, Ananda Theertha Suresh +1
Reasoning in large language models is often discussed as a single capability, but some of its gains may stem from simpler underlying operations. We examine two such primitives, recall and state-tracking, through five controlled task families centered on state-based recall, and compare matched transformer and hybrid architectures with and without reasoning augmentation. Across the suite, reasoning-augmented variants substantially outperform instruction-only variants, often by large margins. This pattern is consistent with the State over Tokens view: externalized reasoning traces help because they carry the intermediate state forward in token space. By contrast, hybrid inductive bias does not yield a uniform advantage in accuracy once reasoning tokens are available. When architectural differences do appear, they follow task structure: the hybrid Think model is more robust on strictly sequential chained updates, whereas the transformer Think model is more robust on flat multi-hop retrieval. We therefore cast the main contribution of this study as a descriptive account of what drives performance on state-based recall tasks: reasoning-token augmentation appears to be the dominant factor, while hybrid advantages are narrower, task-dependent, and potentially more about inference efficiency than overall capability. We also release the codebase and data required to reproduce these results.
Shivam Rawat, Lucie Flek, Florian Mai +1
Lamarr Institute for Machine Learning and Artificial Intelligence · Rheinische Friedrich-Wilhelms-Universität Bonn