Complex reasoning and agentic applications increasingly rely on long-context inference, where growing KV caches increase both memory usage and decoding overhead. Hybrid models reduce these costs by combining Softmax Attention with Gated DeltaNet (GDN) or Kimi Delta Attention (KDA), which maintain fixed-size recurrent states. These states are commonly stored in FP32 and consume substantial GPU memory, while their updates are limited by memory bandwidth. Quantization can reduce both storage footprint and memory traffic, but we find that uniform INT8 and FP8 degrade complex reasoning accuracy, while INT4 and NVFP4 collapse it to near zero. To our knowledge, this is the first study of post-training recurrent-state quantization for GDN and KDA. Our analysis reveals that outliers in GDN and KDA states are concentrated in particular key channels and value dimensions. Learned decay influences how much quantization error is retained. We find that largely the same GDN heads and KDA key channels exhibit slow decay across tasks. Based on these insights, we propose DAMP, which jointly considers quantization error and decay-based error retention to select high-risk key channels offline. Under a fixed storage budget, it retains these channels in FP16 and stores the remainder in INT8. We evaluate DAMP on Qwen3.6-35B, Kimi-Linear-48B and Kimi-K3 across six reasoning and code generation benchmarks. At 9.9 bits per state value, DAMP maintains average accuracy close to FP32. In SGLang, DAMP reduces recurrent-state storage by 69.1%, accelerates the recurrent-state update kernel by up to 2.59x , and lowers full-model time per output token by up to 19.0%.
Figures & tables
Figure 1: Recurrent-state update cost and accuracy–storage trade-off on Qwen3.6-35B. (a) Decode latency breakdown at batch size 256. (b) AIME 2026 accuracy versus effective state-storage bits.
Figure 2: State structure, quantization error, and decay. (a–b) Outliers across key channels and value dimensions in GDN and KDA. (c) INT8 quantization error remains concentrated in a few key channels after value-dimension scaling and reordering. (d–e) KDA decay varies across steps and channels. (f) Channel rankings by effective decay factor remain stable across tasks.
Figure 3: Overview of Damp . Offline calibration selects high-precision key channels. A fused kernel updates the packed FP16/INT8 state, combining low-precision reconstruction, the recurrence, and requantization.
Math Reasoning
General Reasoning
Code
Method
Avg. bits
AIME 2026
HMMT
IMO-Ans
GPQA-D
MMLU-Pro
LCB-v6
Qwen3.6-35B-A3B
FP32
32
85.46
57.57
50.29
81.97
84.66
86.95
FP16
16
84.58
55.82
49.42
82.13
84.61
86.80
BF16
16
79.71
54.97
43.71
81.66
84.65
84.10
FP8 (E4M3)
9.0
29.27
14.06
10.23
76.26
77.66
15.02
Table 1: Main results on Qwen3.6-35B-A3B and Kimi-Linear-48B-A3B-Instruct. “Avg. bits” reports effective state-storage cost per element, including scales and zero points. SR denotes scaling and reordering along the value dimension. Damp results are highlighted in bold .
Figure 4: Decoding efficiency of Damp -INT8, FP32, and BF16/FP16 state storage. (a–c) Recurrent-update latency per layer; (d–f) TPOT reduction relative to FP32. Columns show Qwen3.6-35B-A3B, Kimi-Linear-48B-A3B, and Kimi-K3, respectively.
Method
Avg. bits
AIME 2026
HMMT
FP32
32
96.88
97.92
BF16
16
97.29
96.78
FP8
9.0
59.83
40.91
INT8
9.0
89.17
82.58
INT8+SR
9.0
93.96
91.67
INT4
6.0
4.38
7.58
Table 2: Kimi-K3 accuracy (%).
State storage
TTFT (s) ↓
Mean cache hit rate (%) ↑
Mean
P95
FP32
38.427
72.471
47.20
BF16
35.646
69.500
56.96
Damp INT8
30.464
61.559
65.29
Table 3: Multi-turn TTFT and cache hit rates on Kimi-K3.
Table 4: Matched-budget KDA selector ablation at 9.9 bits per state value.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Qwen3.6-35B
Kimi-Linear-48B
Kimi-K3
AIME 2026
128
128
16
HMMT February 2026
64
64
8
IMO-AnswerBench
16
16
–
GPQA-Diamond
32
32
–
MMLU-Pro
16
16
–
LiveCodeBench-v6
8
8
–
Appendix
Table 5: Number of generations per problem for each model and benchmark.
State format
4K
8K
16K
32K
64K
128K
Qwen3.6-35B-A3B (GDN)
FP32 state
96.96
96.94
96.65
96.56
96.32
96.13
FP16 state
96.98
96.94
96.65
96.56
96.32
96.11
INT8+SR
96.93
96.89
96.60
96.43
96.28
96.07
Damp INT8
97.00
96.94
96.65
96.56
96.32
96.10
Kimi-Linear-48B-A3B-Instruct (KDA)
Appendix
Table 6: RULER macro accuracy (%) across context lengths.
Arch.
Common
Overlap
Spearman
KDA
14.72
92.0%
0.990
GDN
14.91
93.2%
0.986
Random
2.00
12.5%
–
Appendix
Table 7: Agreement between calibration splits, averaged across layers and heads. Common is the number of shared top-16 channels.
Figure 6: Head-level decay structure in Qwen3.6-35B (GDN). (a) Decay factors across decoding steps for heads at p10, p50, and p90 of aeff . (b) Effective decay factors across all 960 heads; shading shows the p10–p90 range across decoding steps. (c) Effective decay factors on Code and General compared with Math, with Spearman rank correlations.
Selector
AIME 2026
GPQA-D
LCB-v6
Random
66.59
80.98
66.88
State energy
82.19
81.37
85.17
Damp ( C )
83.82
82.05
85.67
Appendix
Table 8: Matched-budget GDN selector ablation at 9.9 bits per state value.