Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based token generation. However, their growing model sizes and high inference costs make efficient deployment challenging: full-sequence denoising repeatedly invokes compute-intensive forward passes, while block-diffusion models additionally introduce a memory-intensive KV-cache. Low-bit weight-activation quantization is therefore attractive, yet existing dLLM post-training quantization methods rely on calibration data despite activation distributions shifting across masking states and denoising steps. We present SoloQ, a calibration-free quantization framework that maps weights and activations into a normalized rotated basis with a predictable marginal distribution, enabling data-independent quantization. SoloQ combines a structured K-RPBH rotation with a lightweight rescaling correction for calibration-free quantization. Its predictable post-rotation distribution supports both distribution-matched codebooks and hardware-native NVFP4. For block-diffusion models, SoloQ further applies commit-time KV-cache quantization to compress persistent states without perturbing the actively denoised block. Across full-sequence dLLMs (LLaDA and Dream) and block-diffusion dLLMs(Fast-dLLM v2 and Nemotron-Labs-Diffusion), SoloQ retains accuracy under 4-bit quantization and outperforms calibration-based baselines on knowledge- and reasoning-intensive benchmarks. With NVFP4, SoloQ reduces peak memory by up to 2.61X and accelerates end-to-end inference by up to 2.24X.
Figures & tables
Figure 1: Profiling and motivation for SoloQ . (a) Memory breakdown at batch size 128 and aggregate linear/attention latency at batch size 1 for Fast-dLLM v2 ( Wu et al., 2026a ) across context lengths on an H200 GPU. (b) Layer-wise variation in down_proj input activation scale across denoising steps, normalized to the first step. (c) Accuracy of PTQ ( Ashkboos et al., 2024 ; Frantar et al., 2022 ) with WinoGrande/MMLU calibration versus calibration-free SoloQ .
Figure 2: Overview of SoloQ . (a) K-RPBH combines permutation, block-Hadamard, and cross-block mixing to transform heterogeneous representations into a predictable space. (b) SoloQ -C and SoloQ -N enable calibration-free weight–activation quantization using distribution-matched codebooks and NVFP4, respectively. (c) For block-diffusion models, SoloQ applies commit-time KV quantization, keeping the denoising block in full precision.
Figure 3: Effect of K-RPBH rotation on dLLM activations. Top row shows Dream-7B layer-1 FFN down projection and bottom row shows Nemotron layer-2 FFN down projection. (a,d) Coordinate marginals before and after rotation. (b,e) Conditional coordinate variance across output blocks, normalized by the Haar target 1/d . (c,f) KL divergence to the Haar reference versus transform complexity. N/A denotes an unsupported transform dimension for QuaRot ( Ashkboos et al., 2024 ) .
Model
Method
Calibration-Free
Truth.
ARC-C
Hella.
Wino.
PIQA
MMLU
C-EVAL
Human.
GSM8K
Avg.
LLaDA-Base-8B
FP16
-
47.45
44.03
54.09
74.82
74.81
64.15
69.99
31.71
69.52
58.95
RTN
✓
40.45
41.83
45.40
64.72
67.95
49.26
57.95
14.02
16.56
44.23
AWQ
✗
40.87
42.92
46.14
66.88
69.43
51.22
58.43
20.10
36.88
48.09
QuaRot+GPTQ
✗
42.53
44.20
49.76
69.85
70.75
55.96
56.32
25.33
44.57
51.03
DLLMQuant+
✗
41.53
43.44
46.51
67.87
70.12
51.72
59.38
22.13
40.66
49.26
DLLMQuant++
✗
43.53
44.18
51.00
71.85
73.94
57.77
61.22
28.92
56.25
54.29
Table 1: W4A4 results on full-sequence diffusion language models. Values are task accuracy (%) and their average. SoloQ -C and SoloQ -N denote the codebook and NVFP4 variants, respectively. Bold and underlined indicate the best and second-best quantized results within each model.
Model
Precision
Method
Cb-free
Human-B
Human-P
MBPP-B
MBPP-P
GSM8K
MATH
IFEval
MMLU
GPQA
Avg.
Fast-dLLM v2-7B
FP16
–
–
65.85
59.76
61.90
52.65
84.15
60.30
61.18
66.54
30.80
60.35
W4A4
RTN
✓
0.00
0.00
0.00
0.00
0.23
0.00
7.58
24.73
27.46
6.67
AWQ
✗
0.00
0.00
0.00
0.00
0.00
0.00
9.06
24.98
25.45
6.61
QuaRot+GPTQ
✗
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
SoloQ -C
✓
58.54
53.05
60.05
50.53
80.74
54.86
59.89
64.73
30.58
57.00
SoloQ -N
✓
59.10
56.10
57.40
48.90
82.64
54.42
55.27
63.43
33.71
56.77
Table 2: W4A4 and W4A4KV4 results on block-diffusion language models. Values are task accuracy (%) and average. SoloQ -C and SoloQ -N denote the codebook and NVFP4 variants. Bold and underlined mark the best and second-best quantized results per model and precision. Cb-free denotes whether the method is calibration-free, and N/A denotes an inapplicable method.
Figure 4: End-to-end efficiency of SoloQ . (a) Peak inference VRAM and (b) normalized generation latency across five dLLMs. SoloQ -C and SoloQ -N are evaluated on H200 and RTX Pro 6000 GPUs, respectively, against their corresponding BF16 baselines.
Table 7Table 8
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
Calibration-Free
Truth.
Wino.
PIQA
MMLU
GSM8K
Avg.
LLaDA-Base
FP16
-
45.30
73.64
74.84
65.80
68.92
65.70
RTN
✓
38.80
61.80
69.26
51.05
35.03
51.19
AWQ
✗
40.87
66.88
69.43
51.22
36.88
53.06
QuaRot
✗
43.35
72.22
73.61
62.04
57.39
61.72
FlatQuant
✗
42.90
72.16
74.16
63.80
57.24
62.05
DLLMQuant+
✗
41.53
67.87
70.12
51.72
40.66
54.38
Appendix
Table 6: Comparison with FAIR-Calib under W4A4 quantization on the five benchmarks shared by both evaluation settings. Results for FP16, RTN, QuaRot, FlatQuant, and FAIR-Calib are taken from FAIR-Calib, while the remaining baselines and SoloQ results follow our evaluation. Values are task accuracy (%) and their average. Bold and underlined indicate the best and second-best quantized results within each model.
Model
Method
Truth.
ARC-C
Hella.
Wino.
PIQA
MMLU
C-EVAL
Human.
GSM8K
Avg.
LLaDA-Base
FP16
47.45
44.03
54.09
74.82
74.81
64.15
69.99
31.71
69.52
58.95
RTN
40.45
41.83
45.40
64.72
67.95
49.26
57.95
14.02
16.56
44.23
AWQ
40.87
42.92
46.14
66.88
69.43
51.22
58.43
20.10
36.88
48.09
QuaRot
42.53
44.20
49.76
69.85
70.75
55.96
56.32
25.33
44.57
51.03
DLLMQuant+
41.53
43.44
46.51
67.87
70.12
51.72
59.38
22.13
40.66
49.26
DLLMQuant++
43.53
44.18
51.00
71.85
73.94
57.77
61.22
28.92
56.25
54.29
Appendix
Table 7: Comparison with uniform INT4 quantization on full-sequence diffusion language models. SoloQ -C, SoloQ -N, and SoloQ -I denote the codebook, NVFP4, and uniform INT4 variants, respectively. All SoloQ variants use the same unfolded K-RPBH topology. Values are task accuracy (%) and their average.
Model
Method
Human-B
Human-P
MBPP-B
MBPP-P
GSM8K
MATH
IFEval
MMLU
GPQA
Avg.
Fast-dLLM v2
FP16
65.85
59.76
61.90
52.65
84.15
60.30
61.18
66.54
30.80
60.35
RTN
0.00
0.00
0.00
0.00
0.23
0.00
7.58
24.73
27.46
6.67
AWQ
0.00
0.00
0.00
0.00
0.00
0.00
9.06
24.98
25.45
6.61
QuaRot
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
SoloQ -C
58.54
53.05
60.05
50.53
80.74
54.86
59.89
64.73
30.58
57.00
SoloQ -N
59.10
56.10
57.40
48.90
82.64
54.42
55.27
63.43
33.71
56.77
Appendix
Table 8: Comparison with uniform INT4 quantization on Fast-dLLM v2. SoloQ -C, SoloQ -N, and SoloQ -I denote the codebook, NVFP4, and uniform INT4 variants, respectively. All SoloQ variants use the same unfolded K-RPBH topology. Values are task accuracy (%) and their average. N/A indicates that the method is not applicable to the architecture.
Model
Method
Topology
Truth.
ARC-C
Hella.
Wino.
PIQA
MMLU
C-EVAL
Human.
GSM8K
Avg.
LLaDA-Base
FP16
–
47.45
44.03
54.09
74.82
74.81
64.15
69.99
31.71
69.52
58.95
RTN
–
40.45
41.83
45.40
64.72
67.95
49.26
57.95
14.02
16.56
44.23
AWQ
–
40.87
42.92
46.14
66.88
69.43
51.22
58.43
20.10
36.88
48.09
QuaRot
–
42.53
44.20
49.76
69.85
70.75
55.96
56.32
25.33
44.57
51.03
DLLMQuant++
–
43.53
44.18
51.00
71.85
73.94
57.77
61.22
28.92
56.25
54.29
STaR-Quant
–
49.12
44.23
52.75
72.92
73.85
62.95
64.56
35.98
57.29
57.07
Appendix
Table 9: Comparison of unfolded and folded SoloQ under W4A4 quantization on full-sequence diffusion language models. Unfold explicitly applies the activation-side rotations, whereas Fold absorbs foldable residual-stream rotations into adjacent weights offline.
Model
Method
Topology
Human-B
Human-P
MBPP-B
MBPP-P
GSM8K
MATH
IFEval
MMLU
GPQA
Avg.
Fast-dLLM v2
FP16
–
65.85
59.76
61.90
52.65
84.15
60.30
61.18
66.54
30.80
60.35
RTN
–
0.00
0.00
0.00
0.00
0.23
0.00
7.58
24.73
27.46
6.67
AWQ
–
0.00
0.00
0.00
0.00
0.00
0.00
9.06
24.98
25.45
6.61
QuaRot
–
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
N/A
SoloQ -C
Unfold
58.54
53.05
60.05
50.53
80.74
54.86
59.89
64.73
30.58
57.00
SoloQ -C
Fold
51.22
47.56
52.38
44.97
77.56
49.38
51.20
56.70
27.46
50.94
Appendix
Table 10: Comparison of unfolded and folded SoloQ under W4A4 quantization on Fast-dLLM v2. Unfold explicitly applies activation-side rotations, whereas Fold absorbs foldable residual-stream rotations into adjacent weights offline. Values are task accuracy (%) and their average. N/A indicates that the method is not applicable to the corresponding architecture.
Figure 6: Attention-sink patterns in block-diffusion language models. (a) Fast-dLLM v2 exhibits a concentrated sink at position 2 across representative layers. (b) Nemotron-Labs-Diffusion exhibits a sink at position 0. Brighter colors indicate larger attention weights.
Model
Precision
Sink FP
GSM8K
IFEval
Fast-dLLM v2
W4A4KV4
✓
80.59
56.75
✗
79.98
56.38
W4A4KV2
✓
55.80
41.22
✗
50.57
39.93
Nemotron-Labs-Diffusion
W4A4KV4
✓
90.98
65.48
✗
91.28
65.06
Appendix
Table 11: Effect of preserving attention-sink KV states in full precision under aggressive KV-cache quantization. Sink FP indicates that the identified attention-sink position is kept in full precision while the remaining KV states are quantized.
Figure 7: (a) Hardware setup on ZynQ Ultrascale+ ZCU104 FPGA board. (b) Micro-architecture of the hardware accelerator. (c) Hardware accelerator design parameters.
Diffusion large language models (DLLMs) have recently emerged as a promising alternative to autoregressive LLMs by generating text through iterative masked denoising with bidirectional context. However, their large model sizes and iterative denoising process introduce substantial memory and computational overhead, motivating post-training quantization for efficient deployment. In this paper, we identify two key challenges for low-bit DLLM quantization: state-dependent activation disparity and temporal error accumulation. Masked and unmasked tokens exhibit different activation distributions within each denoising step, while quantization errors can accumulate across steps during iterative decoding. To address these challenges, we propose STaR-Quant, a state-time consistent PTQ framework for DLLMs. STaR-Quant introduces State-Guided Activation Transformation (SGAT) to assign masked and unmasked tokens to different activation transformation spaces with a unified static weight-side transformation. It further introduces Temporal Attention Compensation (TAC) to correct the quantized attention representation via a lightweight block-diagonal affine mapping. Experiments on representative DLLMs demonstrate that STaR-Quant consistently improves low-bit weight-activation quantization over strong PTQ baselines, while delivering up to 1.69x speedup and 3.14x memory saving over FP16 deployment.
Xin Yan, Aqiang Wang, Zhenglin Wan +2
School of Artificial Intelligence, Beijing Normal University, Beijing, China · Department of Computer Science, National University of Singapore, Singapore · Centre for Frontier AI Research, Agency for Science, Technology and Research (A*STAR), Singapore
Diffusion Large Language Models (dLLMs) refine tokens iteratively but commit them irreversibly, leading to a "stability lag" where early decisions remain fragile even after being written. We reveal that Post-Training Quantization (PTQ) error easily flips these borderline decisions at the write frontier, which are then permanently locked in and amplified. To address this, we propose Frontier-Aware Instability-Reweighted Calibration (FAIR-Calib), a two-stage PTQ framework for dLLMs. Stage I probes a full-precision teacher to estimate a position prior that combines frontier hits and masked-stage reliability. Stage II performs off-policy, layer-wise calibration by minimizing a reweighted hidden-state MSE, effectively prioritizing the protection of fragile frontier states without requiring expensive end-to-end diffusion rollouts. We further theoretically justify our weighted objective as a surrogate for output KL divergence. Empirically, FAIR-Calib consistently outperforms state-of-the-art baselines on LLaDA and Dream (W4A4), significantly reducing frontier decision flips and suppressing post-commit mismatches across diverse benchmarks.
Haoyu Huang, Linlin Yang, Sheng Xu +5
National College for Excellent Engineers, Beihang University, Beijing, China · State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China · Independent researcher +4
Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. Post-Training Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data.