Organizations: Department of Computer Science and Engineering, Seoul National University, Seoul, South Korea · School of Computer and Communication Sciences, EPFL, Lausanne, Switzerland · Meta Reality Labs, Redmond, USA
Post-training quantization (PTQ) is a practical path to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation quality. We propose Q-Drift, a sampler-side correction that aims to preserve the intended sampling marginals through a deterministic drift adjustment motivated by generalized marginal-preserving SDEs. Q-Drift uses the calibrated conditional residual variance of quantization error to determine a correction factor at each step. Our SDXL study shows that calibration with as few as 10 paired full-precision/quantized runs remains effective. The resulting sampler correction is plug-and-play with common samplers, diffusion models, and PTQ methods, while incurring negligible overhead at inference. Across six diverse text-to-image models (spanning DiT and U-Net), three samplers (Euler, flow-matching, DPM-Solver++), and two PTQ methods (SVDQuant, MixDQ), Q-Drift improves FID over the corresponding quantized baseline in all seven settings of our main evaluation, with up to 4.79 FID reduction on PixArt-Sigma (SVDQuant W3A4), while preserving CLIP scores. Code is available at https://github.com/sooyoung-ryu/Q-Drift.
Figures & tables
Model
Method
FID ↓
CLIP ↑
PSNR ↑
LPIPS ↓
SSIM ↑
FLUX.1-dev
FP
20.68
25.80
–
–
–
(SVDQuant W3A4)
Quantized
24.14
24.66
16.13
0.491
0.621
Q-Drift
24.06
24.69
16.15
0.488
0.623
FLUX.1-schnell
FP
19.18
26.55
–
–
–
(SVDQuant W3A4)
Quantized
23.10
25.43
13.75
0.529
0.519
Q-Drift
22.13
25.57
13.75
0.520
0.528
Table 1: Main results on MJHQ-30K under aggressive quantization settings. SVDQuant rows use W3A4 quantization, and the MixDQ row uses W4A8 quantization. Boldface in the FID column indicates the lower value between the quantized baseline and Q-Drift for each setting.
Figure 1: Visual comparisons on SDXL (SVDQuant W3A4). For readability, prompt texts are provided in Appendix F . Red boxes indicate the regions shown in the zoomed panels.
Sampler-side method
FID ↓
CLIP ↑
Quantized baseline
31.73
26.39
PTQD
31.51
26.35
QNCD
32.01
26.37
D 2 -DPM, deterministic
74.28
25.65
D 2 -DPM, stochastic
82.87
25.32
Q-Drift
30.69
26.40
Table 2: Comparison with prior sampler-side corrections on SDXL (SVDQuant W3A4, Euler).
Calibration / variant
FID ↓
CLIP ↑
Quantized baseline (no correction)
31.73
26.39
Q-Drift, 1K-prompt reference
30.69
26.40
K=10 , max ∑i∣Δci∣
30.85
26.36
K=10 , min ∑iΔci
30.67
26.36
K=10 , max ∑iΔci
30.52
26.37
scalar unconditional variance
31.95
26.35
Table 3: Calibration size and design ablations on SDXL (SVDQuant W3A4).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method (mechanism)
Statistic estimated
How it enters the sampler
Samplers
PTQD (CNC)
correlation coeff. k
divides the output by 1+k
any
PTQD (BC)
channel-wise mean of Δϵ
subtracted from the output
any
PTQD (VSC)
residual variance σq2
shrinks the injected noise σt
stochastic
TAC-Diff. (NER)
channel-wise coeff. Kt
rescales the output
any
TAC-Diff. (IBC)
input bias
corrects the sampler input xt
any
QNCD (inter)
per-sample noise estimate
subtracted from the output
any
Appendix
Table 4: Mechanisms of sampler-side corrections for quantized diffusion. Methods that combine several mechanisms are listed one row per mechanism. QNCD’s intra-step feature smoothing is omitted, as it acts inside the network rather than on the sampler. “stochastic” marks mechanisms that require σt>0 ; “studied” refers to the deterministic Euler, flow-matching, and DPM-Solver++ samplers evaluated here.
Model
Method
FID ↓
CLIP ↑
PSNR ↑
LPIPS ↓
SSIM ↑
FLUX.1-dev
FP
20.68
25.80
–
–
–
(SVDQuant W4A4)
Quantized
20.49
25.68
22.95
0.194
0.830
Q-Drift
20.55
25.68
22.98
0.193
0.830
FLUX.1-schnell
FP
19.18
26.55
–
–
–
(SVDQuant W4A4)
Quantized
18.62
26.41
18.44
0.251
0.728
Q-Drift
18.69
26.44
18.42
0.251
0.727
Appendix
Table 5: Additional results on MJHQ-30K (milder quantization settings). Rows labeled W4A4 use W4A4 quantization for SVDQuant; ViDiT-Q rows use W8A8 quantization.
Figure 2: Empirical validation of marginal and joint Gaussianity on SDXL (SVDQuant W3A4). Top: Histograms of ϵ^θ(t) and Δϵθ(t) at a latent coordinate, overlaid with fitted Gaussian curves. Bottom: The corresponding 2D joint densities with fitted Gaussian ellipses. All 1D and 2D visualizations are shown at the same latent coordinate, (c,h,w)=(0,64,64) .
Figure 3: Empirical validation of the diagonal-covariance simplification. For each selected timestep, we compare the distribution of absolute correlations from 10,000 random off-diagonal pairs, with each correlation estimated over 5,000 calibration samples, against a shuffled baseline. The shuffled baseline is obtained by randomly permuting a variable in each pair. In each panel, the inset reports the mean, median (med), and 95th percentile (p95) of the absolute-correlation distribution, listed as actual vs shuffled. Top: off-diagonal entries of Σϵ^ϵ^(t) . Middle: off-diagonal entries of ΣΔΔ(t) . Bottom: off-diagonal cross-block entries of Σϵ^Δ(t) for i=j .
Model
Method
FID ↓
Δ FID 95% CI
KID ↓
FLUX.1-dev
FP
20.68
–
3.44
(SVDQuant W3A4)
Quantized
24.14
–
5.67
Q-Drift
24.06
[ − 0.30, 0.09]
5.59
FLUX.1-schnell
FP
19.18
–
2.66
(SVDQuant W3A4)
Quantized
23.10
–
5.03
Q-Drift
22.13
[ − 1.18, − 0.84]
4.52
Appendix
Table 6: Distributional metrics and statistical uncertainty on MJHQ-30K, computed on the same generated image sets used in Table 1 . KID is in ×103 scale and brackets denote 95% paired bootstrap intervals. A Δ interval lying entirely below zero indicates an improvement supported by the sample.
Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challenging in resource-constrained scenarios, such as interactive applications and personal single-GPU usage. Post-training quantization (PTQ) offers a practical solution by compressing pretrained models without expensive retraining. However, existing PTQ methods still suffer from severe quality degradation under extremely low-bit settings. In this paper, we identify channel ordering as an important but underexplored factor in per-group quantization. In this setting, each contiguous group shares one quantization scale. When channels with very different statistics are placed in the same group, the scale can be dominated by outliers and cause large quantization errors. Based on this observation, we propose PermuQuant, a simple and effective PTQ framework for low-bit diffusion models. PermuQuant sorts channels by a joint second-moment criterion before per-group quantization, placing channels with similar activation and weight statistics into the same group. It further uses a calibration-based acceptance rule to apply reordering only when the selected permutation reduces quantization error on calibration data. The selected permutations are absorbed into adjacent modules or applied to weights offline, avoiding explicit runtime permutation operations. Extensive experiments on multiple large diffusion models show that PermuQuant consistently reduces quantization error and outperforms existing PTQ baselines. On FLUX.1-dev with an RTX 5090, PermuQuant achieves up to a 1.7× single step speedup and reduces the DiT memory footprint by 3.5× under W4A4 NVFP4 quantization. Code will be available at https://github.com/yscheng04/PermuQuant.
Yongsen Cheng, Kai Liu, Kaiwen Tao +5
Shanghai Jiao Tong University · Huawei Noah’s Ark Lab
Diffusion models generate conditional samples by progressively denoising Gaussian noise, yet the denoising trajectory can stall at visually plausible but low-quality outcomes with conditional misalignment or structural artifacts. We interpret this behavior as local optima in a surrogate quality landscape: Once early denoising commits to a suboptimal global structure, later steps mainly sharpen details and seldom correct the underlying mistake. While existing inference-time approaches explore alternative diffusion states via re-noising with fixed strength or direction, they exhibit limited capacity to escape steep quality plateaus. We propose Controlled Random Zigzag Sampling (Ctrl-Z Sampling),a scalable sampling strategy that detects plateaus in quality landscape via a surrogate score, and allocates exploration only when a plateau is detected. Upon detection, Ctrl-Z Sampling rolls back to noisier states, samples a set of alternative continuations, and updates the trajectory when a candidate improves the score, otherwise escalating the exploration depth to escape the current plateau. The proposed method is model-agnostic and broadly compatible with existing diffusion frameworks. Experiments show that Ctrl-Z Sampling consistently improves generation quality over other inference-time scaling samplers across different NFE budgets, offering a scalable compute-quality trade-off. Code available at: https://github.com/ShunqiM/Ctrl-Z-Sampling.
Shunqi Mao, Wei Guo, Chaoyi Zhang +3
School of Computer Science, The University of Sydney, Australia
Diffusion Large Language Models (dLLMs) refine tokens iteratively but commit them irreversibly, leading to a "stability lag" where early decisions remain fragile even after being written. We reveal that Post-Training Quantization (PTQ) error easily flips these borderline decisions at the write frontier, which are then permanently locked in and amplified. To address this, we propose Frontier-Aware Instability-Reweighted Calibration (FAIR-Calib), a two-stage PTQ framework for dLLMs. Stage I probes a full-precision teacher to estimate a position prior that combines frontier hits and masked-stage reliability. Stage II performs off-policy, layer-wise calibration by minimizing a reweighted hidden-state MSE, effectively prioritizing the protection of fragile frontier states without requiring expensive end-to-end diffusion rollouts. We further theoretically justify our weighted objective as a surrogate for output KL divergence. Empirically, FAIR-Calib consistently outperforms state-of-the-art baselines on LLaDA and Dream (W4A4), significantly reducing frontier decision flips and suppressing post-commit mismatches across diverse benchmarks.
Haoyu Huang, Linlin Yang, Sheng Xu +5
National College for Excellent Engineers, Beihang University, Beijing, China · State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China · Independent researcher +4