cs.LGSep 27, 2026

SketchSSM: Write to the Full State, Read from a Compact Sketch

Authors: Omin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer, Sehoon Kim, Jae W. Lee

Organizations: Seoul National University · UC Berkeley · KAIST

Abstract

Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent-state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires a full-state read even though the state remains unchanged between state updates. We observe that low-rank state-weighted query approximation accurately preserves state-read outputs. Although future queries are unknown, the basis vectors used to approximate them can be fixed offline. Based on this observation, we introduce SketchSSM, which preserves full-state updates while approximating reads. At each state update, SketchSSM reads the full state once to precompute outputs for these basis vectors, storing them in a compact sketch. Each subsequent decode step combines the sketch vectors with query-dependent coefficients to reconstruct the output without a full-state read. Across four Mamba-2-, GDN-, and KDA-based models, SketchSSM at a mean sketch rank of 8 reduces state-access traffic by approximately 10×\times while matching the average accuracy of the FP32 full-state baseline across four decode benchmarks, and preserves recall on four RULER retrieval tasks. At this rank on one NVIDIA B300, linear-attention kernel speedups over the Standard vLLM baseline reach 7.30×\times, 5.02×\times, and 5.24×\times for Mamba-2, GDN, and KDA, respectively, with up to 2.77×\times higher decode throughput on Nemotron 3 Super.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators

    May 7, 2026Anupama Sridhar, Alexander JohansenKoopman Operator

  2. DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling

    Aug 3, 2026Yixiao Qian, Song Chen, Pengkai Wang +3Recurrent StateKimi Delta Attention

  3. SpecLA: Efficient Speculative Decoding for Linear-Attention Models

    Jul 18, 2026Zhibin Wang, Xuying Han, Zhaohua Yang +5Speculative DecodingKimi Delta Attention