cs.CLSep 29, 2026

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Authors: Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, +1 more

Organizations: Zhejiang University · Tsinghua University · NLPR & MAIS, Institute of Automation, CAS · City University of Hong Kong · Harvard University · The Chinese University of Hong Kong

Abstract

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

    Sep 29, 2026Yi Pan, Haocheng Xi, Kan Zhu +10Large Language Model QuantizationKimi Delta Attention

  2. DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

    Aug 27, 2026Tao Zhang, Jianchao Tan, Pingwei Sun +6Recurrent StateGated Deltanet

  3. Q-Delta: Beyond Key-Value Associative State Evolution

    Jun 7, 2026Sumin Park, Seojin Kim, Noseong ParkKimi Delta AttentionRecurrent State