cs.LGJun 9, 2026

Express Language Modeling

Authors: Albert GongAnnabelle Michael CarrellRaaz DwivediLester Mackey

Organizations: Cornell Tech · University of Cambridge · Microsoft Research

Abstract

We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the state-of-the-art Thinformer approximation, Express improves upon the best known causal attention guarantees, delivering log3/2(n)/s\log^{3/2}(n)/s approximation error with only O(s)O(s) memory and O(s2log2(n))O(s^2 \log^2(n)) compression overhead for a sequence of length nn. We pair these developments with an efficient I/O-aware Triton implementation, demonstrate substantial speedups over FlashAttention 2, and use Express to overcome four resource bottlenecks in the language modeling pipeline: long-context prefill, KV cache compression, long-form memory-constrained decoding, and long-form compute-constrained decoding.

Explore similar work

CardsList