cs.LGSep 30, 2026

Low-Discrepancy Dither for Quantized Recurrent State Caches

Authors: Snigdha Chandan Khilar

Organizations: Independent Researcher

Abstract

Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither, which needs no random numbers, consistently brings the quantized model closer to the full-precision one than stochastic rounding, across pure and hybrid models, storage formats, and long decoding horizons, at no extra cost. Round-to-nearest behaves differently: because it discards small updates, its error keeps growing, so it can look best in short evaluations yet falls far behind over long generations. A discrepancy analysis explains this ordering, and we document implementation pitfalls that silently remove the benefit.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

    Sep 29, 2026Bingchen Yao, Haobo Xu, Haokun Lin +6Recurrent StateKimi Delta Attention

  2. DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

    Aug 27, 2026Tao Zhang, Jianchao Tan, Pingwei Sun +6Recurrent StateGated Deltanet

  3. LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

    Sep 29, 2026Yi Pan, Haocheng Xi, Kan Zhu +10Large Language Model QuantizationKimi Delta Attention