cs.CLOct 7, 2026

Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

Authors: Xiaoran Liu, Ziwei He, Xipeng Qiu

Organizations: Fudan University · Shanghai Innovation Institute · OpenMOSS Team

Abstract

The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16×\times training-free length extrapolation while maintaining 100% accuracy on NIAH-SK1 in 64k context length.

Figures & tables

Appendix figures & tables33 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking the Role of Efficient Attention in Hybrid Architectures

    Jun 13, 2026Ziqing Qiao, Yinuo Xu, Chaojun Xiao +6Efficient Long-Context InferenceDynamic Attention

  2. Learning When to Attend: Conditional Memory Access for Long-Context LLMs

    Mar 18, 2026Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4Efficient Long-Context InferenceTime-To-First-Token