cs.LGOct 8, 2026

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

Authors: Qitong Wang, Xinwei Niu, Mingluo Su, Shanwei Zhao, Shiai Zhu, Huan Wang

Organizations: Westlake University · Ant Group

Abstract

The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding. Nevertheless, typical methods in this line compute the Hessian using pre-collected natural sequences, whereas the model is fed self-generated tokens during decoding, creating a distribution shift between the two sequences. The Hessian calculated on the natural sequence is different from that calculated on the generated sequence. We observe that this discrepancy causes the activation distribution during generation to deviate from that used for pruning, further hurting the pruned model performance. Moreover, most existing LLM pruning methods that bring actual speedup primarily target the sparse matrix-matrix (SpMM) multiplication, providing limited support for the sparse matrix-vector (SpMV) operations, which dominate decoding. To solve these problems, we introduce SparseDecoding, a principled decoding-aware pruning framework tailored for accurate and efficient LLM decoding. Specifically, at the algorithmic axis, SparseDecoding constructs calibration matrices from layer-wise activations collected during the dense-model autoregressive generation, excluding prefill, thereby aligning the pruning objective with the decoding activations. At the system axis, we develop an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal. Substantial empirical results on representative LLMs (Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B / 32B) demonstrate that our method consistently outperforms standard fixed-text calibration on the long-form generation benchmarks while achieving up to 1.48x end-to-end wall-clock decoding speedup on A100 GPUs.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Unified Static-Dynamic Pruning for Efficient LLM Inference

    Jul 24, 2026Jinhyeok Kim, Yejoon Lee, Jaeyoung DoEfficient Neural Network InferenceLLM Pruning

  2. SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference

    Jun 9, 2026Jaeseong Lee, Seung-won Hwang, Samyam RajbhandariLLM PruningHigh-Performance Computing

  3. Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning

    Sep 16, 2026Dunyao Xue, Chengshuo Du, Zhengbo Wang +2Language Model DecodingDiverse Text Generation