cs.CLSep 28, 2026

MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution

Authors: Prasoon Dev, Anirudh Sankar, Vasudeva Varma

Organizations: Language Technologies Research Center International Institute of Information Technology Hyderabad Hyderabad, Telangana, India

Abstract

Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.

Figures & tables

Explore similar work

CardsList
  1. Dynamic Linear Attention

    Jun 9, 2026Xin Wang, Hui Shen, Boyuan Zheng +7Kimi Delta AttentionEfficient Long-Context Inference

  2. Liquid Gated Attention

    Aug 31, 2026Yiheng Jiang, Yuanbo Xu, Yongjian YangTemporal AttentionDiscontinuous Gating