Organizations: Zhejiang University · Tsinghua University · NLPR & MAIS, Institute of Automation, CAS · City University of Hong Kong · Harvard University · The Chinese University of Hong Kong
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.
Figures & tables
Figure 1: Why recurrent state needs structured compression. (a) Recurrent state memory grows with concurrent requests and can exceed model weight memory. STEPQuant substantially reduces memory cost at 6bit. (b) Heads with longer gate half-lives tend to accumulate larger state errors under uniform INT6 quantization across 2304 Qwen heads. (c) Mean accuracy across seven tasks for Qwen (solid) and KDA (dashed). STEPQuant achieves 6.93× and 5.03× compression of Qwen recurrent states at 4 and 6 bits, respectively, with negligible degradation in mean accuracy.
Figure 2: Spatial structure of recurrent states. (a) Key rows are ranked by readout impact and divided into eight groups. Quantizing one group at a time to INT4 generally causes larger perplexity increases for higher-impact groups. (b) A representative Qwen state exhibits obvious outliers in both key rows and value columns. (c) Channel RMS relative to the median across 2048 decoding steps. The outliers remain prominent throughout decoding.
State
LCB v6
EvalPlus
AIME 26
MATH-500
HMMT
GPQA-D
IFBench
Avg.
Qwen3.8-27B
FP32
85.31
84.87
87.71
97.60
74.24
80.81
53.67
80.60
INT8
72.23
80.26
78.54
97.00
58.71
75.25
41.00
71.86
INT6
30.43
69.00
33.96
87.40
16.67
51.52
26.33
45.04
INT4
7.87
31.55
0.00
34.00
0.00
4.04
11.67
12.73
STEPQuant@6bit
85.42
85.54
87.24
97.35
73.30
80.87
54.42
80.59
Table 1: Long-generation reasoning accuracy with BF16 weights (%).
State
LCB v6
EvalPlus
AIME 26
MATH-500
HMMT
GPQA-D
IFBench
Avg.
Qwen3.8-27B
FP32
85.23
85.06
82.92
97.20
71.78
81.06
52.00
79.32
STEPQuant@6bit
85.33
85.42
82.76
97.25
71.02
80.93
52.17
79.27
STEPQuant@4bit
83.98
84.69
85.42
96.40
70.83
81.06
50.50
78.98
Kimi-Linear-48B-A3B-Instruct
FP32
52.80
74.72
59.79
93.80
41.86
66.92
22.75
58.95
Table 3: Reasoning performance of STEPQuant with 4-bit AWQ-quantized weights.
Variant
AIME
GPQA
LCB
Avg.
FP32
87.71
80.81
85.31
84.61
INT6
33.96
51.52
30.43
38.63
Q-Mamba@6bit
73.13
76.77
77.97
75.95
Spatial only
79.79
80.68
80.09
80.19
Temporal w/o pivots
60.42
71.97
69.00
67.13
Temporal only
74.58
79.67
72.64
75.63
Table 4: Component ablation on three long-reasoning tasks using Qwen3.8-27 BF16 weights.
Figure 3: Generation length and serving efficiency of STEPQuant. (a–b) Mean generated tokens across seven tasks, weighting tasks equally. Blue bars show FP32 and uniform INT8/6/4. Orange bars show STEPQuant@6bit/@4bit (marked @6/@4). Lengths include thinking, incorrect answers, and capped outputs. (c) Total serving memory of Qwen with W4 weights across different batch sizes. (d) Normalized recurrent-state memory and updating time of Qwen at batch size 512.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B1: Gate lifetime and key-row impact of recurrent state. (a–b) Cumulative INT6 squared error (red) and unit count (blue), ordered by increasing gate half-life within each layer or head. The longest-lived quarter accounts for 52.5% of Qwen error and 78.8% of KDA error across all 2,304 heads and 81,920 channels. (c) KDA half-life versus accumulated INT6 error for 79,864 channels after excluding near-zero errors; color encodes error. (d) Held-out equal-norm readout-error contrast across row-impact octiles within fixed-precision groups, using independent WikiText ranking.
Figure B2: Dual-axis state geometry across Qwen and KDA. Four native-order 128×128 states (Qwen, Qwen, KDA, KDA). Top: entries below 10× the matrix-median absolute magnitude are white; larger entries use graded greens. Marginal traces show row and column RMS relative to their axis medians. Bottom: matching surfaces use the same style as the Qwen surface in Figure 2 (b). Floor traces show the same normalized RMS profiles, rescaled for display; heights are capped at 80× . Each state has seven or eight rows and columns above 5× their axis-median RMS.
Figure B3: Global lifetime rankings remain similar across text sources. Qwen (2,304 heads) and KDA (81,920 key channels) are ranked separately, from shortest to longest gate half-life. Each point compares its global WikiText rank percentile with its C4 (orange-red) or LiveCodeBench (sky blue) percentile. The thin black diagonal is WikiText compared with itself ( y=x ). Spearman coefficients use the complete population in each panel.
Property
Qwen3.8-27B
Kimi-Linear-48B-A3B
Recurrent layers
48
20
State heads per recurrent layer
48
32
Allocation unit
entire head
key row
Number of allocation units
2,304
81,920
Integer candidates (@4 / @6)
{2,4,6,8} / {4,6,8}
{2,4,6,8} / {4,6,8}
Optimizer
multiple-choice DP
Lagrangian allocation
Appendix
Table C1: Model-specific STEPQuant configuration. FP16 pivot and scale-metadata costs are accounted for separately from the nominal bit budget.
1. Reconstruct the previous state from integer codes and scales, or from FP16 values for pivot units.
2. Compute X=DtSt−1+βtkt(vt⊤−kt⊤DtSt−1) .
3. Emit yt=X⊤qt before requantization and continue subsequent model computation.
4. Derive row factors using Key-Row-Aware Dual-axis Fitting (Section 4.3 ). Qwen’s two-bit path shares row factors within value groups.
5. For KDA, jointly fit one shared column-scale vector with squared row-impact factors over all non-pivot integer rows within each head. Qwen fits each head at its assigned precision. Its two-bit path uses signed magnitude levels.
6. Pack integer codes and scales. Store pivot values in FP16. Keep the old representation valid until its readers finish.
Appendix
Table C2: One persistent STEPQuant decode step. The same logical workflow supports four-bit and six-bit budgets.
Model
Weights
STEPQuant@6bit
STEPQuant@4bit
Qwen
BF16
6.361
4.621
Qwen
W4A16
6.361
4.621
KDA
BF16
6.300
4.300
KDA
W4A16
6.300
4.300
Appendix
Table D3: Compact recurrent-state representation (bits/value). Integer codes, FP16 scales, and pivot replacement are included. Qwen@4 uses the packed byte count above; the other entries follow Equations 22 and 23 .
State
LCB v6
EvalPlus
AIME 26
MATH-500
HMMT
GPQA-D
IFBench
Avg.
Qwen3.8-27B
FP32
5.64
0.62
9.73
1.67
18.15
5.10
4.71
6.52
INT8
9.74
0.87
15.44
2.32
29.95
6.04
9.02
10.48
INT6
19.63
1.43
25.62
5.68
33.24
11.31
16.72
16.23
INT4
29.51
2.17
53.28
36.30
50.74
48.62
53.53
39.16
STEPQuant@6bit
4.44
0.64
8.51
1.62
20.11
5.12
5.89
6.62
Appendix
Table E4: Mean generated tokens with BF16 weights (thousands).
State
LCB v6
EvalPlus
AIME 26
MATH-500
HMMT
GPQA-D
IFBench
Avg.
Qwen3.8-27B
FP32
8.99
0.91
10.87
1.66
18.00
4.97
5.11
7.22
STEPQuant@6bit
8.77
0.89
11.03
1.67
21.37
5.02
5.19
7.70
STEPQuant@4bit
10.29
0.95
13.56
1.80
23.88
5.25
10.80
9.50
Kimi-Linear-48B-A3B-Instruct
FP32
9.27
1.35
24.62
4.07
34.20
8.92
6.49
12.70
Appendix
Table E5: Mean generated tokens with W4A16 weights. Values are thousands of output tokens over all evaluated samples, with the same definition and task ordering as Table E4 .
Variant
AIME
GPQA
LCB
Avg.
FP32
68.33
69.70
54.52
64.18
INT6
22.71
57.83
42.37
40.97
Q-Mamba@6bit
38.54
57.07
44.55
46.72
Spatial only
63.54
66.67
52.80
61.00
Temporal w/o pivots
41.46
59.85
48.87
50.06
Temporal only
61.67
66.16
48.91
58.91
Appendix
Table E6: Component ablation on Kimi-Linear-48B-A3B-Instruct. Accuracy (%) with BF16 weights. Avg. denotes the average across three long-generation benchmarks.
Figure F4: Serving memory and state-update time of STEPQuant on KDA.
Model
Batch
FP32
STEPQuant@6bit
Gain
Qwen
32
1,656
1,727
+4.32%
Qwen
64
2,980
3,122
+4.77%
Qwen
128
4,157
4,485
+7.87%
Qwen
256
5,655
6,360
+12.47%
Qwen
512
6,040
7,280
+20.53%
KDA
32
5,248
5,356
+2.06%
Appendix
Table F7: Decode throughput. Units are tokens/s. Gains use the unrounded throughputs.
Figure F5: Full-model decode throughput.
Method
Avg. bits
Retention (%)
DAMP
9.9
100.99
STEPQuant@6bit
6.3
100.51
STEPQuant@4bit
4.3
92.91
INT8 †
9.0
83.54
INT4 †
5.0
6.66
NVFP4 †
4.5
6.74
Appendix
Table G8: Accuracy retention on three shared KDA benchmarks: AIME 2026, HMMT Feb 2026, and LiveCodeBench v6. Retention is normalized to each study’s FP32 baseline.
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70× at the kernel level and 1.47× for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
Complex reasoning and agentic applications increasingly rely on long-context inference, where growing KV caches increase both memory usage and decoding overhead. Hybrid models reduce these costs by combining Softmax Attention with Gated DeltaNet (GDN) or Kimi Delta Attention (KDA), which maintain fixed-size recurrent states. These states are commonly stored in FP32 and consume substantial GPU memory, while their updates are limited by memory bandwidth. Quantization can reduce both storage footprint and memory traffic, but we find that uniform INT8 and FP8 degrade complex reasoning accuracy, while INT4 and NVFP4 collapse it to near zero. To our knowledge, this is the first study of post-training recurrent-state quantization for GDN and KDA. Our analysis reveals that outliers in GDN and KDA states are concentrated in particular key channels and value dimensions. Learned decay influences how much quantization error is retained. We find that largely the same GDN heads and KDA key channels exhibit slow decay across tasks. Based on these insights, we propose DAMP, which jointly considers quantization error and decay-based error retention to select high-risk key channels offline. Under a fixed storage budget, it retains these channels in FP16 and stores the remainder in INT8. We evaluate DAMP on Qwen3.6-35B, Kimi-Linear-48B and Kimi-K3 across six reasoning and code generation benchmarks. At 9.9 bits per state value, DAMP maintains average accuracy close to FP32. In SGLang, DAMP reduces recurrent-state storage by 69.1%, accelerates the recurrent-state update kernel by up to 2.59x , and lowers full-model time per output token by up to 19.0%.
Tao Zhang, Jianchao Tan, Pingwei Sun +6
South China University of Technology · Meituan · East China Normal University
Linear attention reformulates sequence modeling as recurrent state evolution, enabling efficient linear-time inference. Under the key-value associative paradigm, existing approaches restrict the role of the query to the readout operation, decoupling it from state evolution. We show that query-conditioned state readout induces a structured value prediction over accumulated memory that complements key-based retrieval. Based on this insight, we propose Q-Delta, a query-aware delta rule that integrates mixed key-query prediction errors into state evolution, enabling jointly corrective dynamics while preserving delta-rule efficiency. We establish stability guarantees for the resulting dynamics and derive a hardware-efficient chunkwise-parallel formulation with a custom Triton implementation. Empirical results demonstrate stable optimization, competitive throughput, and consistent improvements over strong baselines on language modeling and long-context retrieval tasks.
Sumin Park, Seojin Kim, Noseong Park
Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea.