Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70× at the kernel level and 1.47× for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
Figures & tables
Figure 1: Recurrent-state access is a major cost of LLM serving. (a) Decode-time breakdown for GLM-5.3-Flash-NVFP4 on the B200 GPU at batch size 256. The share spent on linear attention stays roughly constant as context length grows. (b) GPU memory footprint for Qwen3.5-9B with prefix caching, assuming one linear attention state per 1,024 cached tokens.
Figure 2: Per-token versus per-window quantization. Per-token quantization re-quantizes the full state after every update. Per-window quantization holds the quantized boundary state S^0 fixed, buffers the p updates of the window in higher precision, and reconstructs the state from them; the full state is quantized only once per window, into S^p .
Figure 3: State MSE over 64K decoded tokens on PG-19 (Qwen3.5-9B); -w quantizes once per window.
Figure 4: Per-window state reconstruction with Compensator Tokens. The FP32 state is split into a low-bit quantized residual, a few higher-precision Compensator Tokens that capture its dominant large-magnitude structure, and the buffered real-token updates of the window. The Compensator Tokens absorb dominant outliers, flattening the residual’s row ℓ2 norms and making it easier to quantize.
Figure 5
Qwen3.5-9B
Qwen3.5-35B-A3B
Kimi-Linear-48B-A3B
Method
AIME
GPQA
LCB
MMLU
AIME
GPQA
LCB
MMLU
AIME
GPQA
LCB
MMLU
Avg.
FP32
87.9
81.3
64.1
83.3
91.5
84.7
75.6
85.9
67.5
70.3
41.4
72.4
75.5
BF16
72.1
66.2
49.6
81.0
85.8
79.3
67.2
85.2
64.3
68.1
41.0
64.0
68.7
8-bit methods
Ours
87.9
81.8
64.1
83.8
91.0
83.9
76.1
85.8
68.3
69.8
41.3
72.1
75.5
FP8
14.6
34.3
21.4
42.6
29.6
39.9
26.7
56.4
25.6
46.6
16.0
57.0
34.2
Table 1: Downstream performance (%, higher is better). Best scores in each column within the 8-bit, 6-bit, and 4-bit groups are in bold.
Figure 7: Kernel throughput of one linear attention layer.
Figure 8: Decode-step throughput in vLLM at context length 4K. On the RTX PRO 6000, the largest batch size is 256 where 512 does not fit, and Qwen3.5-35B-A3B uses 2K instead of 4K at this batch size.
Figure 9: End-to-end inference throughput on different datasets.
Method
AIME
LCB
Kernel
FP32 per-step
87.9
64.1
1.00 ×
BF16 per-step
72.1
49.6
1.64 ×
+ Per-window
87.8
62.9
1.71 ×
INT8 per-step
7.1
9.2
2.43 ×
+ Per-window
82.4
60.6
2.64 ×
+ Comp. Tokens
86.6
61.5
2.56 ×
Table 2: Ablation of individual LeapQuant components on Qwen3.5-9B.
p
AIME
LCB
Kernel
4
86.2
62.8
1.59 ×
8
87.1
63.2
2.13 ×
16
87.9
64.1
2.52 ×
32
87.9
64.2
2.20 ×
Table 3: Ablations of window length p and Compensator Token count r on Qwen3.5-9B.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Dt
k~t
bt
Record size
Linear Attention
I
kt
0
dk+dv
RetNet
γI
kt
0
dk+dv
Mamba2
αtI
kt
0
1+dk+dv
GLA / RWKV6 / HGRN2
\Diag(αt)
kt
0
2dk+dv
DeltaNet
I
βtkt
kt
dk+dv
Gated DeltaNet
αtI
βtkt
αtkt
1+dk+dv
Appendix
Table 4: Instances of equation 1 . DeltaProduct applies n delta-rule sub-steps per token, each of which is an instance of equation 1 . The last column gives the number of values in one buffered update (di,k~i,ui) .
Qwen3.5-9B
Qwen3.5-35B-A3B
Kimi-Linear-48B-A3B
Method
GSM
AIME
GPQA
LCB
MMLU
GSM
AIME
GPQA
LCB
MMLU
GSM
AIME
GPQA
LCB
MMLU
FP32
96.1
87.9
81.3
64.1
83.3
96.7
91.5
84.7
75.6
85.9
92.1
67.5
70.3
41.4
72.4
BF16
94.9
72.1
66.2
49.6
81.0
96.6
85.8
79.3
67.2
85.2
92.1
64.3
68.1
41.0
64.0
8-bit methods
Ours
96.3
87.9
81.8
64.1
83.8
96.5
91.0
83.9
76.1
85.8
92.0
68.3
69.8
41.3
72.1
FP8 per-tensor
79.2
14.6
34.3
21.4
42.6
86.7
29.6
39.9
26.7
56.4
91.8
25.6
46.6
16.0
57.0
Appendix
Table 5: Full downstream performance (%, higher is better). Best scores in each column within each bit width are in bold.
Figure 10: Decode-step throughput in vLLM on the B200 at context lengths 1K–8K. Each bar stacks the FP32 throughput (lighter) and the gain of LeapQuant (darker).
Figure 11: Decode-step throughput of the two sparse-attention models on the B200 at batch size 256 and long contexts.
Model
Temp.
Top- p
Top- k
Presence
Thinking
Max output
Qwen3.5-9B
1.0
0.95
20
1.5
on
81,920
Qwen3.5-35B-A3B
1.0
0.95
20
1.5
on
81,920
Kimi-Linear-48B-A3B
1.0
1.0
–
0
–
65,536
Appendix
Table 6: Generation settings of the downstream evaluation.