Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path's coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.
Figures & tables
Figure 1: Overview of QuantMLA. Our analysis reveals pronounced RoPE-path error amplification under matched cache reconstruction error. Guided by this analysis, we introduce QuantMLA, which learns path-specific transformations with function-aligned objectives and fuses them entirely offline for accurate and efficient low-bit MLA inference.
Figure 2: Layer-wise functional response under matched cache reconstruction error. Curves show the content- and RoPE-path responses Rb across layers for six MLA models. The adjacent distributions show the corresponding layer-wise RoPE/content response ratios; diamonds and whiskers denote medians and interquartile ranges, and dashed lines mark equal response.
Figure 3: QuantMLA learning and execution pipeline. Content and RoPE transformations are optimized with attention-output and positional QK reconstruction, respectively, and fused offline. The native low-bit MLA backend stores packed cache representations while serving protected sink and recent tokens from BF16 buffers. QDQ denotes simulated quantize–dequantize.
Precision
Method
General-purpose
Information extraction
Avg.
CS
MMLU
GSM8K
FDA
SWDE
SQuAD
DeepSeek-V2-Lite
BF16
—
69.76
57.90
36.92
78.13
89.20
57.10
66.80
C4R4
RTN
67.17
51.90
16.98
65.06
84.88
55.50
61.02
SmoothQuant †
68.23
54.02
24.56
70.69
87.85
56.07
63.44
QuaRot †
69.01
55.66
30.71
72.14
88.12
55.06
64.67
Table 1: General-purpose and information-extraction performance. Scores (%, ↑ ). CS averages the five commonsense benchmarks; Avg. assigns equal weight to all ten tasks. Bold marks the best quantized score within each precision, including ties. † : our MLA adaptations.
Precision
Method
Reasoning
Code
Avg.
GPQA-
MMLU
GSM8K
MATH
AIME25
Human
LiveCode
Diamond
500
Eval
Bench
LongCat-Flash-Lite
BF16
—
46.46
80.77
71.72
57.40
70.00
73.78
49.67
64.26
C4R4
RTN
42.42
76.38
68.01
48.00
40.00
39.02
32.70
49.50
SmoothQuant †
42.42
78.49
69.75
53.20
53.33
60.37
42.75
57.19
Table 2: Reasoning and code performance. Scores (%, ↑ ). Avg. assigns equal weight to all seven benchmarks. Bold marks the best quantized score within each precision. † : our MLA adaptations.
Figure 4: Retrieval as the cached history grows. DeepSeek-V2-Lite on RULER S-NIAH-1/2/3 at C4R4, with BF16 as reference. Each point evaluates 500 examples; method markers follow Table 1 .
Configuration
MMLU (%) ↑
BF16
57.90
Unprotected C4R4 baselines
RTN
51.90
SmoothQuant †
54.02
QuaRot †
55.66
QuantMLA cumulative components
Table 3: Baseline comparison and component ablation at C4R4. Scores on DeepSeek-V2-Lite MMLU; parenthesized values denote incremental gains within the QuantMLA sequence.
Figure 5: KV-cache storage and inference efficiency of C4R4 with sink4/local128. (a) Persistent KV-cache allocation, including BF16 buffers and metadata. (b) Attention latency at batch sizes 8 and 32, normalized to paired measurements with an unquantized BF16 cache. (c) Whole-job output throughput on eight GPUs under the same native vLLM scheduling policy; labels report preemption counts. Panel (c) compares one C4R4 run against a historical BF16 baseline.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Content response
RoPE response
Ratio
DeepSeek-V2-Lite
0.620
3.925
6.331
DeepSeek-V3-Base
0.452
1.751
3.871
DeepSeek-R1
0.494
1.784
3.611
Kimi-K2-Instruct
1.594
3.095
1.941
LongCat-Flash-Lite
0.594
4.077
6.864
GLM-4.7-Flash
1.026
2.462
2.400
Appendix
Table 4: Model-level responses at matched cache NMSE. Ratios use unrounded source aggregates; displayed responses are rounded.
Figure 6: Layer-wise atlas of dual-path output distortion. Upper: content (C) and RoPE (R) output NMSE at cache NMSE 0.01 . Lower: RoPE/content ratios, with white marking equal distortion. Logarithmic color scales are shared across models; gray pads positions beyond each model’s analyzed depth.
Figure 7: Predicting and testing functional amplification. Layer-wise gains, 72 validation points, and equal-energy interventions separate operator sensitivity from quantization-error geometry at cache NMSE 10−3 .
Model
Spearman
Median absolute log error
DeepSeek-V2-Lite
0.9923
0.1054
GLM-4.7-Flash
0.9614
0.0727
Predictor
DeepSeek-V2-Lite log-ratio MAE
GLM-4.7-Flash log-ratio MAE
Cache error only
1.5657
0.8790
Operator trace only
0.9129
1.4276
Operator + error allocation
0.1133
0.1173
Appendix
Table 5: Predictive validation of the operator–error model. Each model contributes 36 validation points. MAE denotes mean absolute error.
Error intervention
Content
RoPE
Independent sign randomization
1.048
1.115
Whole-token permutation
5.918
1.495
Isotropic direction
6.007
1.380
Appendix
Table 6: Equal-energy interventions isolate quantization-error allocation. Geometric-mean gain ratios relative to the original QDQ-induced errors use 32 repetitions per control across four DeepSeek-V2-Lite layer-20 configurations. Unity indicates unchanged gain.
Item
Setting
Calibration corpus
WikiText-2 training split
Sequence length / batch size
2048 / 4
Train / validation bank
128 / 32 sequences per seed
Optimization
300 steps per path
Optimizer
AdamW, weight decay 0
Schedule
10% warmup, cosine decay
Appendix
Table 7: Calibration and transformation-learning settings. Learning uses disjoint training and validation banks; downstream comparisons use seed 0.
Figure 8: Native mixed-precision MLA execution. The fused writer generates a packed C4R4 copy for every incoming token and retains the pre-quantization BF16 values of protected sink and recent tokens in dedicated buffers. Attention uses pre-quantization BF16 values for protected tokens and dequantized INT4 values otherwise. When a non-sink token leaves the recent window, execution switches to its existing packed copy without additional quantization. Reconstructed content tiles support both matching and latent aggregation, while all valid tokens share one global softmax normalization.
Table 9: Attention latency at fixed context lengths on GPU ( μ s). C4R4 uses the mixed-precision cache policy (sink4/local128). Each entry averages 120 samples from two independent runs; Δ=(tC4R4/tBF16−1)×100% . BF16 uses native partitioning. Here 1K=1024 and 1M=1,048,576 tokens.
Metric
BF16 KV
C4R4 KV
Usable cache slots
657,856
2,360,832
Elapsed time (s)
1,958.392
378.929
Output throughput (tokens/s)
2.614
13.512
Preemption events
32
0
Recomputed tokens
3,996,255
0
Appendix
Table 10: Serving under native vLLM scheduling on eight GPUs. Five simultaneous requests each use 128K input and 1K output tokens. C4R4 includes sink4/local128 protection.
Query-key (QK) normalization stabilizes attention by controlling the scale of queries and keys before the dot product, but is not immediately compatible with Multi-head Latent Attention (MLA). MLA achieves efficient decoding by caching low-dimensional latent states instead of full keys, whereas post-projection QK RMSNorm appears to require the fully projected key for every cached token. We show this apparent incompatibility is an implementation artifact, not an architectural constraint. RMSNorm decomposes into a static affine weight and a dynamic scalar RMS statistic. The static key-side weight can be absorbed into the MLA query-side projection; the dynamic key statistic reduces to one inverse-RMS scalar per token and KV group. The resulting formulation is exactly equivalent to explicit post-projection QK RMSNorm in exact arithmetic and preserves MLA's latent decode path. In our 400M runs trained for up to 100B tokens, QK-Normed MLA achieves lower training loss and better downstream accuracy than QK clipping, while H800 decode benchmarks show less than 2% latency overhead up to 256k context. These results make QK normalization a practical stabilization option for MLA models without requiring full-key caching.
Yizhou Han, Yao Zhao, Jun Zhou +2
School of Data Science, The Chinese University of Hong Kong, Shenzhen.
Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decoding memory traffic. Yet most capable open checkpoints use multi-head or grouped-query attention (MHA/GQA), so conversion is needed to obtain MLA's cache efficiency without retraining from scratch. Speculative decoding offers complementary acceleration, but its speedup depends on agreement between draft proposals and target verification. We find that direct MHA/GQA-to-MLA conversion can sharply reduce this agreement: low-rank factorization and RoPE handling introduce attention-function errors that may be tolerable for standalone generation but substantially lower draft-token acceptance. We therefore formulate MLA draft construction as functional reconstruction rather than cache compression. Our end-to-end (E2E) method optimizes each converted MLA attention module to reproduce the post-output-projection response of its original MHA/GQA counterpart on calibration hidden states. This converter-agnostic post-conversion procedure preserves the converted cache and inference graph and requires neither verifier logits nor verifier supervision. We evaluate 192 model-converter-backend-method-task configurations spanning four Llama/Qwen draft-target pairs, TransMLA and MHA2MLA, HF and vLLM, and four 200-prompt tasks. With a 0.5-percentage-point reporting tolerance, Functional Reconstruction materially improves acceptance in 37 of 64 matched task cells, leaves 26 practically unchanged, and materially decreases one. Code and evaluation artifacts are available at https://github.com/swyhahaha/FunctionalMLA.
Weiye Shi, Fanxu Meng, Muhan Zhang
Institute for Artificial Intelligence, Peking University
Multi-head Latent Attention (MLA), the attention used in DeepSeek-V2/V3, jointly compresses keys and values into a low-rank latent and matches the H100 roofline almost perfectly. Its trained weights, however, expose only one decoding path - an absorbed MQA form - which ties efficient inference to H100-class compute-bandwidth ratios, forfeits tensor parallelism along the head axis, and yields no Multi-Token Prediction (MTP) gain on commodity inference GPUs such as the export-restricted H20. We propose Group-Query Latent Attention (GQLA), a minimal modification of MLA whose trained weights expose two algebraically equivalent decoding paths over the same parameters: an MQA-absorb path identical to MLA's, and a GQA path with a per-group expanded cache. The runtime picks the path that matches the target hardware - no retraining, no custom kernels - so a single set of GQLA weights pins the rooflines of both H100 (MQA-absorb, s_q=1) and H20 (GQA + MTP, s_q=2), while supporting up to 8-way zero-redundancy tensor parallelism on the GQA path. To avoid pretraining from scratch we extend TransMLA into TransGQLA, which converts a pretrained GQA checkpoint into a GQLA model; on LLaMA-3-8B it compresses the per-token KV cache to 28.125% of the GQA baseline on the MQA-absorb path while structurally preserving GQA-level traffic on the per-group path.
Fanxu Meng
Institute for Artificial Intelligence, Peking University