Organizations: The University of Sydney · The Hong Kong University of Science and Technology · The Chinese University of Hong Kong · Shanghai AI Laboratory
Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry to guide offline calibration. Isolated block--step interventions estimate propagation risk, which prioritizes sensitive trajectory states during row-radius selection. With these radii fixed, response-subspace correction uses neighboring-code edits to reduce residual components along dominant activation directions. Both stages preserve the original 4-bit weight representation. Controlled interventions show that short-horizon propagated error predicts final latent error more reliably than immediate block-output error, supporting calibration beyond local reconstruction objectives. Evaluations on Wan models, Self Forcing, and MiniMax-H3 demonstrate improvements in key consistency and dense-reference metrics while remaining competitive on other attributes across model scales and generation paradigms.
Figures & tables
Figure 1: Propagation diagnostics from 80 isolated 4-bit interventions on Wan 2.1-14B (10 blocks × eight denoising steps), each followed by dense continuation. (a) Local block-output RMSE is weakly associated with final latent error. (b) Four-step propagated error EH predicts final damage after ranking blocks within each timestep. (c) Final damage depends strongly on trajectory position. Colors encode block depth and markers encode injection timestep.
Figure 2: Overview of PulseQuant. Isolated 4-bit pulses reveal how local perturbations propagate through denoising (left). The resulting risk weights guide per-row radius selection through a mean–tail output-error objective (center). With the selected radii fixed, legal neighboring-code edits reduce residual components along high-energy activation directions (right). The calibrated weights are exported as packed 4-bit indices and per-row scales, with no online profiling or correction.
Model
Method
W/A
VBench attributes (%)
Dense-reference fidelity
IQ ↑
AQ ↑
MS ↑
BC ↑
SC ↑
Scene ↑
OC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Wan 2.1-1.3B
Dense BF16
16/16
67.71
66.06
98.29
96.19
93.97
42.44
26.15
ref.
ref.
ref.
SmoothQuant ( Xiao et al., 2023 )
4/6
59.83
58.12
98.12
94.77
91.14
30.96
24.20
14.78
0.464
0.499
SVDQuant ( Li et al., 2024 )
4/6
63.68
63.57
97.83
95.20
92.62
35.61
25.42
15.36
0.513
0.433
ViDiT-Q ( Zhao et al., 2024 )
4/6
63.60
63.68
97.78
95.03
92.70
38.08
25.73
14.86
0.503
0.447
DVD-Quant ( Li et al., 2026c )
4/6
63.21
64.04
98.62
95.23
94.38
41.86
25.60
14.96
0.529
0.430
Table 1: Quantitative evaluations on VBench (T2V). VBench attributes are percentages; PSNR, SSIM, and LPIPS compare matched quantized and Dense BF16 videos. We bold the best results and underline the second-best results .
Model
Method
W/A
VBench attributes (%)
Dense-reference fidelity
IQ ↑
AQ ↑
MS ↑
DD ↑
BC ↑
SC ↑
Scene ↑
OC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Wan 2.1-T2V (#Param:14B)
Dense BF16
16/16
67.14
67.12
98.48
66.67
96.69
94.36
39.46
26.46
ref.
ref.
ref.
ViDiT-Q ( Zhao et al., 2024 )
4/6
67.25
66.23
98.58
56.94
96.30
94.10
41.72
26.31
15.89
0.545
0.378
OrbitQuant ( Lee et al., 2026 )
4/6
67.18
67.14
98.60
62.50
96.44
95.01
37.86
26.22
16.93
0.585
0.325
PulseQuant (ours)
4/6
67.22
67.34
98.73
61.11
96.56
95.21
42.81
26.52
16.90
0.579
0.331
ViDiT-Q ( Zhao et al., 2024 )
4/4
59.41
58.36
98.13
30.56
95.00
91.59
25.29
23.60
13.27
0.415
0.567
Table 2: Quantitative evaluations of scale and architecture transfer on VBench (T2V).
Figure 3: MiniMax-H3 W4A4 comparison at frames 5, 60, and 100 under the same condition.
Configuration
IQ ↑
AQ ↑
OC ↑
PSNR ↑
LPIPS ↓
Uniform PTQ
56.90
57.29
23.43
14.00
0.549
+ Rotation
63.20
62.39
24.70
14.97
0.443
+ Risk
64.38
64.97
25.73
15.59
0.414
+ Subspace corr.
64.64
64.24
25.74
15.61
0.412
PulseQuant
65.83
64.93
26.18
16.08
0.385
Table 3: Component ablation on Wan 2.1-T2V-1.3B W4A6 with λ=0.5 and ρ=0.25 .
Configuration
IQ ↑
AQ ↑
OC ↑
PSNR ↑
LPIPS ↓
Uniform PTQ
56.90
57.29
23.43
14.00
0.549
+ Rotation
63.20
62.39
24.70
14.97
0.443
+ Risk
64.38
64.97
25.73
15.59
0.414
+ Subspace corr.
64.64
64.24
25.74
15.61
0.412
PulseQuant
65.83
64.93
26.18
16.08
0.385
Table 3: Component ablation on Wan 2.1-T2V-1.3B W4A6 with λ=0.5 and ρ=0.25 .
Direction
IQ ↑
AQ ↑
DD ↑
Scene ↑
OC ↑
None
62.17
62.90
51.39
35.17
25.50
Random
61.62
62.86
48.61
37.72
25.38
Fixed Walsh
61.96
62.24
51.39
34.52
25.08
Calibrated (ours)
62.70
62.46
55.56
37.78
25.64
Table 4: Ablation results of correction-direction ablation on Wan 2.1-T2V-1.3B W4A4 with λ=0.5 and ρ=0.25 .
Figure 4: Sensitivity and profiling choices. (a–b) Risk-objective robustness on Wan 2.1-T2V-1.3B W4A4. (c–d) Propagation-horizon stability and anchor-grid agreement on Wan 2.1-T2V-14B.
Figure 5: Response-axis adaptation lowers subspace error with code edits and weight MSE in Wan-14B.
Figure 6: RTX 5080 W4A4 efficiency: (a) DiT memory; (b,c) VC/SageAttention2 per-step latency and speedup relative to BF16.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Setting
Default
Calibration data
Number of prompts
3
Calibration data
Prompt composition
Complementary coverage
Propagation profile
Horizon H
4 steps
Propagation profile
Anchor grid
10 blocks ×8 timesteps
Risk objective
Tail weight λ
0.75
Risk objective
Tail fraction ρ
0.5
Appendix
Table 5: Default PulseQuant calibration configuration. Model-specific exceptions are reported with the evaluation protocol.
Shared 16-value codebook, entry, and coordinate index
QA,ba
Activation quantizer and activation bit width
ri(0),ri,rˉi
Initial row norm (also denoted ri ) and calibrated radius
Ri,Kr,η
Radius candidates, grid size, and relative half-span
Appendix
Table 6: Notation used in the method and implementation details.
Components
VBench attributes (%)
Dense-reference fidelity
Configuration
Rotation
Risk
Subsp. corr.
IQ ↑
AQ ↑
MS ↑
BC ↑
OC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
T-HP ↓
Uniform PTQ
56.90
57.29
97.38
95.01
23.43
14.00
0.421
0.549
0.161
Rotation only
✓
63.20
62.39
98.09
95.01
24.70
14.97
0.512
0.443
0.155
Risk only
✓
✓
64.38
64.97
98.18
95.58
25.73
15.59
0.529
0.414
0.150
Subspace corr. only
✓
✓
64.64
64.24
98.26
95.69
25.74
15.61
0.536
0.412
0.142
PulseQuant
✓
✓
✓
65.83
64.93
98.47
95.30
26.18
16.08
0.546
0.385
0.144
Appendix
Table 7: Complete component ablation on Wan 2.1-T2V-1.3B W4A6 under matched 300-layer coverage.
Figure 7: Layer-level activation and error diagnostics on Wan 2.1-14B W4A4 over one matched 50-step calibration trajectory. (a,b) Per-channel peak activation magnitudes before and after scaling and RPBH for the most outlier-heavy 128-channel group; planes and stars mark the 99th percentile and maximum, respectively. (c) Four-step propagation gains measured on 10 blocks ×8 timestep anchors; intermediate timesteps are interpolated only for visualization. (d,e) RMS output-response residuals relative to BF16 for naive W4A4 and PulseQuant, evaluated on 128 sampled output channels with a shared scale; bars report relative RMSE. (f) Pointwise residual reduction, where positive values favor PulseQuant. At this audited layer, relative RMSE decreases from 9.0% to 5.3%, and 96% of channel–timestep locations improve.
λ
ρ
VBench attributes (%)
Dense-reference fidelity
N-Jerk ↓
IQ ↑
AQ ↑
MS ↑
DD ↑
BC ↑
SC ↑
Scene ↑
OC ↑
TF ↑
PSNR ↑
SSIM ↑
LPIPS ↓
ρ sweep with λ=0.75
0.75
0.125
61.97
62.03
97.78
51.39
94.62
92.02
38.15
25.23
98.60
15.45
0.506
0.442
1.429
0.75
0.250
62.50
62.40
97.80
52.78
94.73
91.89
35.61
25.41
98.59
15.43
0.503
0.443
1.429
0.75
0.375
62.78
62.37
97.79
56.94
94.91
91.76
35.03
25.24
98.57
15.38
0.503
0.444
1.429
0.75
0.500
62.70
62.46
97.83
55.56
94.84
91.97
36.70
25.44
98.57
15.42
0.503
0.443
1.427
Appendix
Table 8: Propagation-risk sensitivity on Wan 2.1-T2V-1.3B W4A4. Gray rows mark the default (λ,ρ)=(0.75,0.5) .
Calibration set
ID
Calibration prompt
Complementary coverage
P1
A ceramic teapot rotating slowly on a wooden table in soft studio light
P2
A glass marble rolling across a linen cloth while the camera tracks beside it
P3
A small sailboat crossing a calm lake at dawn with gentle camera movement
Similar static
P1
A ceramic teapot resting on a wooden table under soft studio light, locked-off camera
P2
A polished red apple resting on a linen-covered table under soft studio light, locked-off camera
P3
A porcelain vase resting on a wooden shelf under soft studio light, locked-off camera
Appendix
Table 9: Exact three-prompt calibration sets for the MiniMax-H3 composition study; prompts are reproduced verbatim.
Calibration set
# Prompts
VBench attributes (%)
Dense-reference fidelity
IQ ↑
AQ ↑
MS ↑
DD ↑
BC ↑
SC ↑
Scene ↑
OC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Dense BF16
–
68.57
65.38
99.01
75.00
96.08
92.55
47.60
26.64
ref.
ref.
ref.
Prompt count: shared ordered prefix
Prefix-1
1
66.39
65.55
98.73
68.06
95.33
92.42
46.51
25.97
14.26
0.424
0.509
Prefix-3
3
66.83
65.43
98.73
70.83
94.97
92.23
46.88
26.24
14.17
0.424
0.514
Prefix-6
6
66.74
65.75
98.74
68.06
95.29
92.65
40.84
26.15
14.22
0.426
0.512
Appendix
Table 10: MiniMax-H3 W4A4 calibration-set sensitivity on the VBench.
Figure 8: Wan 2.1-1.3B W4A4 comparison for the Eiffel Tower under lightning.
Figure 9: Wan 2.1-1.3B W4A4 comparison for an animated kitten at a garden fountain.
Figure 10: Wan 2.1-1.3B W4A4 comparison for a woman walking along a beach at sunset.
Figure 11: Wan 2.1-1.3B W4A6 comparison for a brown bear in a forest clearing.
Figure 12: MiniMax-H3 W4A6 comparison for a porcelain cup on a carved wooden coaster.
Figure 13: MiniMax-H3 W4A6 comparison for a woman relaxing by the beach at sunset.
Figure 14: MiniMax-H3 W4A4 comparison for a zebra walking across a sunlit savanna.
Figure 15: MiniMax-H3 W4A4 comparison for yellow flowers swaying beside a wooden fence.
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive. Post-training quantization (PTQ) is the natural remedy, yet DiT activations shift across timesteps, prompts, and guidance branches, forcing prior methods to re-fit calibration data for every new checkpoint or modality. We present OrbitQuant, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis. In this basis, a randomized permuted block-Hadamard (RPBH) rotation concentrates each coordinate around one fixed, known marginal regardless of the input, so a single Lloyd-Max codebook serves all timesteps, prompts, and layers of a given input dimension. We extend the same quantizer to weight rows offline, absorbing the rotation into the weights so that it cancels inside each linear layer and only a forward rotation on the activations remains at runtime. The same recipe transfers from image to video with no per-modality tuning. Across FLUX.1, Z-Image-Turbo, Wan 2.1, and CogVideoX, it sets the state of the art for PTQ at several low-bit settings. It also pushes PTQ of image diffusion transformers to W2A4 with usable generation quality.
Donghyun Lee, Jitesh Chavan, Duy Nguyen +5
Cantina Labs · University of Southern California · *Work done during an internship at Cantina Labs. +1
Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent. The standard fix applies an invertible linear transform to the activations and its inverse to the weights before quantizing both. Normalization layers between blocks force this transform to run online at every denoising step, making its inference computation cost the binding design constraint. Existing options trade quantization quality for inference cost: per-channel scaling (SmoothQuant) is computationally cheap but impacts the magnitude of the channels, which can harm quantization accuracy; fixed Hadamard transforms yield better quantization accuracy but require large block sizes that incur a high online cost; learned full-d invertible transforms calibrate best but entail a prohibitive dense d×d matrix multiplication (GEMM) per layer per step. We propose KroQuant, a PTQ method that applies a learned Kronecker-structured invertible transform to each 32-element block of the activation, storing less than half the parameters of per-channel scaling. The block-local structure runs as small tensor-core GEMMs, and on an MI350 GPU the KroQuant quantizer kernel is up to 14% faster than the SmoothQuant kernel. Offline LoRaQ weight calibration then absorbs the residual per-weight quantization error. On PixArt-Σ, SANA, and FLUX.1-schnell at W4A4 (MXFP4e2), KroQuant produces outputs closer to the FP reference than SVDQuant and LoRaQ on MJHQ-30K and SDCI, while preserving or improving image quality.
We present a post-training quantization (PTQ) approach for Wan2.1-T2V-14B, a 14-billion-parameter text-to-video diffusion transformer, targeting the W8A8 HiFloat8 (HiF8) format on Ascend 910B NPUs. A central challenge in quantizing video DiT models is the heterogeneous activation distribution across transformer blocks: boundary blocks (the first and last few blocks) exhibit fundamentally different statistical properties from middle blocks, making uniform quantization ineffective. We conduct a systematic per-block activation analysis across all 40 WanAttentionBlocks and use the findings to motivate a boundary-protection strategy that retains the first two and last three blocks in BF16 while quantizing the remaining 35 blocks with W8A8 HiF8. The proposed PTQ method matches or marginally exceeds the BF16 baseline on all five VBench dimensions evaluated, indicating no measurable accuracy loss within the 5-prompt evaluation set. An ablation study over four protection configurations confirms that full boundary protection yields the highest average VBench score, validating the data-driven block selection. We additionally investigate quantization-aware training (QAT) as a complementary fine-tuning stage and analyze the conditions under which it fails to outperform plain PTQ on single-card hardware.
Yiming Zhao
University of Chinese Academy of Sciences Beijing, China