Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path's coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.
Figures & tables
Figure 1: Overview of QuantMLA. Our analysis reveals pronounced RoPE-path error amplification under matched cache reconstruction error. Guided by this analysis, we introduce QuantMLA, which learns path-specific transformations with function-aligned objectives and fuses them entirely offline for accurate and efficient low-bit MLA inference.
Figure 2: Layer-wise functional response under matched cache reconstruction error. Curves show the content- and RoPE-path responses Rb across layers for six MLA models. The adjacent distributions show the corresponding layer-wise RoPE/content response ratios; diamonds and whiskers denote medians and interquartile ranges, and dashed lines mark equal response.
Figure 3: QuantMLA learning and execution pipeline. Content and RoPE transformations are optimized with attention-output and positional QK reconstruction, respectively, and fused offline. The native low-bit MLA backend stores packed cache representations while serving protected sink and recent tokens from BF16 buffers. QDQ denotes simulated quantize–dequantize.
Precision
Method
General-purpose
Information extraction
Avg.
CS
MMLU
GSM8K
FDA
SWDE
SQuAD
DeepSeek-V2-Lite
BF16
—
69.76
57.90
36.92
78.13
89.20
57.10
66.80
C4R4
RTN
67.17
51.90
16.98
65.06
84.88
55.50
61.02
SmoothQuant †
68.23
54.02
24.56
70.69
87.85
56.07
63.44
QuaRot †
69.01
55.66
30.71
72.14
88.12
55.06
64.67
Table 1: General-purpose and information-extraction performance. Scores (%, ↑ ). CS averages the five commonsense benchmarks; Avg. assigns equal weight to all ten tasks. Bold marks the best quantized score within each precision, including ties. † : our MLA adaptations.
Precision
Method
Reasoning
Code
Avg.
GPQA-
MMLU
GSM8K
MATH
AIME25
Human
LiveCode
Diamond
500
Eval
Bench
LongCat-Flash-Lite
BF16
—
46.46
80.77
71.72
57.40
70.00
73.78
49.67
64.26
C4R4
RTN
42.42
76.38
68.01
48.00
40.00
39.02
32.70
49.50
SmoothQuant †
42.42
78.49
69.75
53.20
53.33
60.37
42.75
57.19
Table 2: Reasoning and code performance. Scores (%, ↑ ). Avg. assigns equal weight to all seven benchmarks. Bold marks the best quantized score within each precision. † : our MLA adaptations.
Figure 4: Retrieval as the cached history grows. DeepSeek-V2-Lite on RULER S-NIAH-1/2/3 at C4R4, with BF16 as reference. Each point evaluates 500 examples; method markers follow Table 1 .
Configuration
MMLU (%) ↑
BF16
57.90
Unprotected C4R4 baselines
RTN
51.90
SmoothQuant †
54.02
QuaRot †
55.66
QuantMLA cumulative components
Table 3: Baseline comparison and component ablation at C4R4. Scores on DeepSeek-V2-Lite MMLU; parenthesized values denote incremental gains within the QuantMLA sequence.
Figure 5: KV-cache storage and inference efficiency of C4R4 with sink4/local128. (a) Persistent KV-cache allocation, including BF16 buffers and metadata. (b) Attention latency at batch sizes 8 and 32, normalized to paired measurements with an unquantized BF16 cache. (c) Whole-job output throughput on eight GPUs under the same native vLLM scheduling policy; labels report preemption counts. Panel (c) compares one C4R4 run against a historical BF16 baseline.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Content response
RoPE response
Ratio
DeepSeek-V2-Lite
0.620
3.925
6.331
DeepSeek-V3-Base
0.452
1.751
3.871
DeepSeek-R1
0.494
1.784
3.611
Kimi-K2-Instruct
1.594
3.095
1.941
LongCat-Flash-Lite
0.594
4.077
6.864
GLM-4.7-Flash
1.026
2.462
2.400
Appendix
Table 4: Model-level responses at matched cache NMSE. Ratios use unrounded source aggregates; displayed responses are rounded.
Figure 6: Layer-wise atlas of dual-path output distortion. Upper: content (C) and RoPE (R) output NMSE at cache NMSE 0.01 . Lower: RoPE/content ratios, with white marking equal distortion. Logarithmic color scales are shared across models; gray pads positions beyond each model’s analyzed depth.
Figure 7: Predicting and testing functional amplification. Layer-wise gains, 72 validation points, and equal-energy interventions separate operator sensitivity from quantization-error geometry at cache NMSE 10−3 .
Model
Spearman
Median absolute log error
DeepSeek-V2-Lite
0.9923
0.1054
GLM-4.7-Flash
0.9614
0.0727
Predictor
DeepSeek-V2-Lite log-ratio MAE
GLM-4.7-Flash log-ratio MAE
Cache error only
1.5657
0.8790
Operator trace only
0.9129
1.4276
Operator + error allocation
0.1133
0.1173
Appendix
Table 5: Predictive validation of the operator–error model. Each model contributes 36 validation points. MAE denotes mean absolute error.
Error intervention
Content
RoPE
Independent sign randomization
1.048
1.115
Whole-token permutation
5.918
1.495
Isotropic direction
6.007
1.380
Appendix
Table 6: Equal-energy interventions isolate quantization-error allocation. Geometric-mean gain ratios relative to the original QDQ-induced errors use 32 repetitions per control across four DeepSeek-V2-Lite layer-20 configurations. Unity indicates unchanged gain.
Item
Setting
Calibration corpus
WikiText-2 training split
Sequence length / batch size
2048 / 4
Train / validation bank
128 / 32 sequences per seed
Optimization
300 steps per path
Optimizer
AdamW, weight decay 0
Schedule
10% warmup, cosine decay
Appendix
Table 7: Calibration and transformation-learning settings. Learning uses disjoint training and validation banks; downstream comparisons use seed 0.
Figure 8: Native mixed-precision MLA execution. The fused writer generates a packed C4R4 copy for every incoming token and retains the pre-quantization BF16 values of protected sink and recent tokens in dedicated buffers. Attention uses pre-quantization BF16 values for protected tokens and dequantized INT4 values otherwise. When a non-sink token leaves the recent window, execution switches to its existing packed copy without additional quantization. Reconstructed content tiles support both matching and latent aggregation, while all valid tokens share one global softmax normalization.
Table 9: Attention latency at fixed context lengths on GPU ( μ s). C4R4 uses the mixed-precision cache policy (sink4/local128). Each entry averages 120 samples from two independent runs; Δ=(tC4R4/tBF16−1)×100% . BF16 uses native partitioning. Here 1K=1024 and 1M=1,048,576 tokens.
Metric
BF16 KV
C4R4 KV
Usable cache slots
657,856
2,360,832
Elapsed time (s)
1,958.392
378.929
Output throughput (tokens/s)
2.614
13.512
Preemption events
32
0
Recomputed tokens
3,996,255
0
Appendix
Table 10: Serving under native vLLM scheduling on eight GPUs. Five simultaneous requests each use 128K input and 1K output tokens. C4R4 includes sink4/local128 protection.