Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70× at the kernel level and 1.47× for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
Figures & tables
Figure 1: Recurrent-state access is a major cost of LLM serving. (a) Decode-time breakdown for GLM-5.3-Flash-NVFP4 on the B200 GPU at batch size 256. The share spent on linear attention stays roughly constant as context length grows. (b) GPU memory footprint for Qwen3.5-9B with prefix caching, assuming one linear attention state per 1,024 cached tokens.
Figure 2: Per-token versus per-window quantization. Per-token quantization re-quantizes the full state after every update. Per-window quantization holds the quantized boundary state S^0 fixed, buffers the p updates of the window in higher precision, and reconstructs the state from them; the full state is quantized only once per window, into S^p .
Figure 3: State MSE over 64K decoded tokens on PG-19 (Qwen3.5-9B); -w quantizes once per window.
Figure 4: Per-window state reconstruction with Compensator Tokens. The FP32 state is split into a low-bit quantized residual, a few higher-precision Compensator Tokens that capture its dominant large-magnitude structure, and the buffered real-token updates of the window. The Compensator Tokens absorb dominant outliers, flattening the residual’s row ℓ2 norms and making it easier to quantize.
Figure 5
Qwen3.5-9B
Qwen3.5-35B-A3B
Kimi-Linear-48B-A3B
Method
AIME
GPQA
LCB
MMLU
AIME
GPQA
LCB
MMLU
AIME
GPQA
LCB
MMLU
Avg.
FP32
87.9
81.3
64.1
83.3
91.5
84.7
75.6
85.9
67.5
70.3
41.4
72.4
75.5
BF16
72.1
66.2
49.6
81.0
85.8
79.3
67.2
85.2
64.3
68.1
41.0
64.0
68.7
8-bit methods
Ours
87.9
81.8
64.1
83.8
91.0
83.9
76.1
85.8
68.3
69.8
41.3
72.1
75.5
FP8
14.6
34.3
21.4
42.6
29.6
39.9
26.7
56.4
25.6
46.6
16.0
57.0
34.2
Table 1: Downstream performance (%, higher is better). Best scores in each column within the 8-bit, 6-bit, and 4-bit groups are in bold.
Figure 7: Kernel throughput of one linear attention layer.
Figure 8: Decode-step throughput in vLLM at context length 4K. On the RTX PRO 6000, the largest batch size is 256 where 512 does not fit, and Qwen3.5-35B-A3B uses 2K instead of 4K at this batch size.
Figure 9: End-to-end inference throughput on different datasets.
Method
AIME
LCB
Kernel
FP32 per-step
87.9
64.1
1.00 ×
BF16 per-step
72.1
49.6
1.64 ×
+ Per-window
87.8
62.9
1.71 ×
INT8 per-step
7.1
9.2
2.43 ×
+ Per-window
82.4
60.6
2.64 ×
+ Comp. Tokens
86.6
61.5
2.56 ×
Table 2: Ablation of individual LeapQuant components on Qwen3.5-9B.
p
AIME
LCB
Kernel
4
86.2
62.8
1.59 ×
8
87.1
63.2
2.13 ×
16
87.9
64.1
2.52 ×
32
87.9
64.2
2.20 ×
Table 3: Ablations of window length p and Compensator Token count r on Qwen3.5-9B.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Dt
k~t
bt
Record size
Linear Attention
I
kt
0
dk+dv
RetNet
γI
kt
0
dk+dv
Mamba2
αtI
kt
0
1+dk+dv
GLA / RWKV6 / HGRN2
\Diag(αt)
kt
0
2dk+dv
DeltaNet
I
βtkt
kt
dk+dv
Gated DeltaNet
αtI
βtkt
αtkt
1+dk+dv
Appendix
Table 4: Instances of equation 1 . DeltaProduct applies n delta-rule sub-steps per token, each of which is an instance of equation 1 . The last column gives the number of values in one buffered update (di,k~i,ui) .
Qwen3.5-9B
Qwen3.5-35B-A3B
Kimi-Linear-48B-A3B
Method
GSM
AIME
GPQA
LCB
MMLU
GSM
AIME
GPQA
LCB
MMLU
GSM
AIME
GPQA
LCB
MMLU
FP32
96.1
87.9
81.3
64.1
83.3
96.7
91.5
84.7
75.6
85.9
92.1
67.5
70.3
41.4
72.4
BF16
94.9
72.1
66.2
49.6
81.0
96.6
85.8
79.3
67.2
85.2
92.1
64.3
68.1
41.0
64.0
8-bit methods
Ours
96.3
87.9
81.8
64.1
83.8
96.5
91.0
83.9
76.1
85.8
92.0
68.3
69.8
41.3
72.1
FP8 per-tensor
79.2
14.6
34.3
21.4
42.6
86.7
29.6
39.9
26.7
56.4
91.8
25.6
46.6
16.0
57.0
Appendix
Table 5: Full downstream performance (%, higher is better). Best scores in each column within each bit width are in bold.
Figure 10: Decode-step throughput in vLLM on the B200 at context lengths 1K–8K. Each bar stacks the FP32 throughput (lighter) and the gain of LeapQuant (darker).
Figure 11: Decode-step throughput of the two sparse-attention models on the B200 at batch size 256 and long contexts.
Model
Temp.
Top- p
Top- k
Presence
Thinking
Max output
Qwen3.5-9B
1.0
0.95
20
1.5
on
81,920
Qwen3.5-35B-A3B
1.0
0.95
20
1.5
on
81,920
Kimi-Linear-48B-A3B
1.0
1.0
–
0
–
65,536
Appendix
Table 6: Generation settings of the downstream evaluation.
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.
Bingchen Yao, Haobo Xu, Haokun Lin +6
Zhejiang University · Tsinghua University · NLPR & MAIS, Institute of Automation, CAS +3
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .
Jiatong Ding, Bingxin Xing, Yu Zhang +9
Shanghai Jiao Tong University · Xi’an Jiaotong University · Xiamen University +1
Complex reasoning and agentic applications increasingly rely on long-context inference, where growing KV caches increase both memory usage and decoding overhead. Hybrid models reduce these costs by combining Softmax Attention with Gated DeltaNet (GDN) or Kimi Delta Attention (KDA), which maintain fixed-size recurrent states. These states are commonly stored in FP32 and consume substantial GPU memory, while their updates are limited by memory bandwidth. Quantization can reduce both storage footprint and memory traffic, but we find that uniform INT8 and FP8 degrade complex reasoning accuracy, while INT4 and NVFP4 collapse it to near zero. To our knowledge, this is the first study of post-training recurrent-state quantization for GDN and KDA. Our analysis reveals that outliers in GDN and KDA states are concentrated in particular key channels and value dimensions. Learned decay influences how much quantization error is retained. We find that largely the same GDN heads and KDA key channels exhibit slow decay across tasks. Based on these insights, we propose DAMP, which jointly considers quantization error and decay-based error retention to select high-risk key channels offline. Under a fixed storage budget, it retains these channels in FP16 and stores the remainder in INT8. We evaluate DAMP on Qwen3.6-35B, Kimi-Linear-48B and Kimi-K3 across six reasoning and code generation benchmarks. At 9.9 bits per state value, DAMP maintains average accuracy close to FP32. In SGLang, DAMP reduces recurrent-state storage by 69.1%, accelerates the recurrent-state update kernel by up to 2.59x , and lowers full-model time per output token by up to 19.0%.
Tao Zhang, Jianchao Tan, Pingwei Sun +6
South China University of Technology · Meituan · East China Normal University