cs.CLSep 27, 2026

GSM: Efficient Language Modeling with Shared Global State

Authors: Yunao Zheng, Bin Wen, Xiaojie Wang, Kaiyu Jiang, Xuanyu Zheng, Changyi Liu, Hongyi Fu, Jianxiong Wang, +7 more

Organizations: Beijing University of Posts and Telecommunications. · Kuaishou Technology.

Abstract

Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing

    Jun 15, 2026Kuzey Torlak, Hüseyin Arda Arslan, Anıl Dervişoğlu +2Efficient Long-Context InferenceHybrid Transformer

  2. RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

    Feb 12, 2026Kaicheng Xiao, Haotian Li, Liran Dong +1Retrieval Layer

  3. DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling

    Aug 3, 2026Yixiao Qian, Song Chen, Pengkai Wang +3Recurrent StateKimi Delta Attention