Long-Context Language Modeling

Latest papers 112

All topics
CardsList
  1. Long Context Pre-Training with Lighthouse Attention

    May 7, 2026Bowen Peng, Subho Ghosh, Jeffrey QuesnelleLanguage Model PretrainingSelf-Attention

  2. Training Transformers for KV Cache Compressibility

    May 7, 2026Yoav Gelberg, Yam Eitan, Michael Bronstein +2Long-Context Language ModelingKV-Cache Compression

  3. Caracal: Causal Architecture via Spectral Mixing

    Apr 30, 2026Bingzheng Gan, Tianyi Zhang, Yusu Li +4Long-Context Language ModelingAutoregressive Language Modeling

  4. Kwai Summary Attention Technical Report

    Apr 27, 2026Chenglong Chu, Guorui Zhou, Guowang Zhang +35Self-AttentionKV Caching

  5. Screening Is Enough

    Apr 1, 2026Ken M. NakanishiSelf-AttentionLong-Context Language Modeling

  6. Learning When to Attend: Conditional Memory Access for Long-Context LLMs

    Mar 18, 2026Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4Self-AttentionLong-Context Language Modeling

  7. RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

    Feb 12, 2026Kaicheng Xiao, Haotian Li, Liran Dong +1Memory-Augmented Neural NetworksLong-Context Language Modeling

  8. Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

    Jan 17, 2026Xingyue Huang, Xueying Ding, Mingxuan Ju +3Self-AttentionLong-Context Language Modeling

  9. TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish

    Dec 28, 2025Melikşah Türker, A. Ebrar Kızıloğlu, Onur Güngör +1Language Model PretrainingLong-Context Language Modeling

  10. Controllably Efficient Language Models

    Nov 7, 2025Jatin Prakash, Aahlad Puli, Rajesh RanganathEfficient Transformer InferenceLong-Context Language Modeling

  11. Concertina: Data-Centric Adaptive Pipeline Parallelism for Efficient Heterogeneous Long-Context LLM Training

    Sep 25, 2025Shiju Wang, Yujie Wang, Fangcheng Fu +6Long-Context Language ModelingEfficient Language Model Training

  12. Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

    Sep 12, 2025Rupert Mitchell, Kristian KerstingLanguage Model PretrainingSoftmax Attention

  13. The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

    Apr 24, 2025Piotr Nawrot, Robert Li, Renjie Huang +3Transformer InferenceLLM Evaluation

  14. LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

    Feb 20, 2025Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev +4Transformer InterpretabilityLong-Context Language Modeling

  15. AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

    Date pendingShaowen Wang, Yuke Zheng, Tansheng Zhu +4Rotary Positional EmbeddingsSelf-Attention