Organizations: Peking University, Melon Group · Alibaba Group · Tsinghua University · Harbin Institute of Technology · Peking University · University of Electronic Science and Technology of China
Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the dominant terminal errors originate in the high-noise structure-generation stage, and terminal-aligned training corrects terminal errors that substantially extended step-local training cannot. This yields a simple staging principle: \emph{first adapt the sparse architecture into a coarse prior, then correct the terminal distribution}. We instantiate the principle as \method, a unified acceleration framework for visual generation that combines a short sparse warm-up, few-step trajectory-mixed distillation, and FP8 quantization with fused kernels. \method sustains 97% attention sparsity with strong visual quality on long-sequence 720P generation across Wan2.1/Wan2.2 backbones and T2V/I2V tasks, and 90% sparsity on Wan2.1-T2V-1.3B-480P. With 3-step CFG-free inference, \method achieves a 265× end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 (220× on H100), and denoises a Wan2.1-T2V-1.3B-480P video in 1.3s.
Figures & tables
Figure 1 : SparkDiffusion teaser. From one pretrained dense video DiT, SparkDiffusion delivers high quality frames across Wan2.1/Wan2.2, T2V/I2V, and 480P/720P with 3-step inference, sustaining 97% attention sparsity on long-sequence 720P models and 90% on the compact Wan2.1-T2V-1.3B-480P.
Figure 2 : End-to-end latency and speedup of SparkDiffusion over Full Attention across Wan video generation models on NVIDIA H100 and RTX 5090 GPUs. Suffixes “90” and “97” denote attention sparsity levels. Speedup is the latency ratio of the full SparkDiffusion stack (few-step distillation, attention sparsity, and FP8) against the dense baseline; NFE conventions follow footnotes 1 .
Figure 3 : Headline speedup. On Wan2.1-T2V-14B-720P, SparkDiffusion composes three multiplicative factors—3-step CFG-free distillation, 97% compensated sparse attention, and fused FP8 kernels—into a measured 265× end-to-end speedup over the dense baseline (footnotes 1 ).
Figure 4 : The high-sparsity trap is rooted in high-noise structure errors. Oracle probe : from the same initial noise, the sparse model’s velocity is replaced by the dense teacher’s velocity inside a short window of sampling steps (“ fix ”; σ in the legend is the flow-matching noise level; Wan2.1-T2V-14B-480P, 50-step sampler, 32 prompts × 4 seeds; error bars: 95% CI). (a) Terminal error to the dense reference output (latent MSE, ↓ ; the teacher is the zero reference and is off the log scale). (b,c) Gains over the uncorrected sparse model ( ↑ ; dashed: dense teacher). At 97% sparsity, fixing the five highest-noise steps removes most of the terminal error, while fixing the five lowest-noise steps leaves it essentially unchanged (beige band: equal-budget gap between the two 5-step fixes); a wider 28-step high-noise window recovers nearly all of it. Both the student and the teacher are sampled with classifier-free guidance, and the intervention replaces the student’s guided velocity with the teacher’s guided velocity, isolating sparsity-induced error from any CFG-removal effect.
Compensation
Term. error ↓
VBench ↑
Rank-64 (default)
0.124
80.92
Full rank ( r=128 )
0.121
81.05
Table 1 : Terminal error (paired latent MSE) at 97% sparsity with rank-64 vs. full-rank ( r=128 ) compensation (Wan2.1-T2V-14B-480P, 10,000 steps).
Figure 5 : Sliced- W2 distance to the data distribution along the denoising trajectory on six toy 2D sequence manifolds. The sparse model is trained until validation loss plateaus, while the distilled student is obtained by trajectory-mixed distillation.
Figure 6 : Video frame comparison on Wan2.1-T2V-14B-480P under attention sparsity levels 80% , 90% , 95% , and 97% .
Figure 7 : Validation loss and raw quality metrics for Wan2.1-14B across sparsity levels. The “Dense pixel MSE” panel reports the terminal error of Sec. 3 after VAE decoding, sharing the pixel-space footing of the other quality panels.
Figure 8 : Qualitative comparison at 97% attention sparsity on Wan2.1-T2V-14B-720P. We compare the dense teacher, the multi-step sparse model, and the few-step sparse model obtained after trajectory-mixed distillation.
Figure 9 : Overview of the SparkDiffusion framework (stages detailed in Sec. 4.1 – 4.3 ).
Figure 10 : Sparse warm-up provides the coarse prior for terminal-aligned distillation. On Wan2.1-T2V-14B-480P, all models are followed by the same trajectory-mixed distillation stage; a short sparse warm-up establishes the usable coarse prior on which distillation builds.
VBench
VBench-2.0
Latency (s)
Model
Method
Total ↑
Total ↑
Creat. ↑
Common. ↑
Control. ↑
Human
Fid. ↑
Physics ↑
Sp.
5090 ↓
H100 ↓
Wan2.1-T2V
1.3B
Full
83.21
56.02
54.73
57.38
34.96
78.71
54.30
0%
182
92
FastWan (VSA)
82.37
54.63
52.66
56.52
32.27
79.69
52.01
90%
2.8
1.2
TurboDiffusion
82.52
54.61
51.72
56.94
31.65
79.55
53.21
90%
2.0
1.0
Table 2 : Comparison of generation quality and efficiency on VBench [ 6 ] and VBench-2.0 [ 37 ] . We evaluate Wan2.1-T2V-1.3B at 480×832 , and Wan2.1-T2V-14B and Wan2.2-T2V-A14B at 720×1280 , all with 81 frames. VBench-2.0 reports the overall score and five capability categories: Creativity, Commonsense, Controllability, Human Fidelity, and Physics. “Sp.” denotes attention sparsity; NFE conventions follow footnotes 1 . All SparkDiffusion rows are measured with the full deployment stack active—W8A8 FP8 quantization with fused kernels—for both quality metrics and latency; baselines are evaluated with their official inference pipelines, including their own quantization settings. Full Attention rows use our strongest dense implementation: BF16 with FlashAttention-2 on RTX 5090 and FlashAttention-3 on H100, and no other acceleration technique. For Wan2.2-T2V-A14B-720P, the RTX 5090 latency includes the overhead of swapping the high-noise and low-noise expert groups in GPU memory; on H100 both expert groups reside in memory simultaneously. Bold marks the best result among accelerated methods within each model block.
V-JEPA 2
VideoMAE V2
Method
Sp.
Cos ↑
ℓ2↑
Cos ↑
ℓ2↑
Full Attention
0%
0.125
27.15
0.0252
2.83
FastWan (VSA)
90%
0.075
21.31
0.0117
2.05
TurboDiffusion
90%
0.078
21.67
0.0125
2.21
SparkDiffusion (Ours)
97%
0.087
23.04
0.0142
2.47
Table 3 : Quantitative diversity comparison on Wan2.1-T2V-14B-720P. Each video is encoded by two frozen video encoders, V-JEPA 2 [ 1 ] and VideoMAE V2 [ 21 ] ; we report the average pairwise cosine and ℓ2 distances among N=5 videos sampled from the same prompt with 5 noise seeds ( 1,000 prompts; protocol follows Shaul et al. [16] ).
Stage-2 objective
VBench ↑
VBench-2.0 ↑
Full Attention (dense, 50-step)
83.69
60.20
PCM only (3-step)
81.94
56.41
DMD only (3-step)
82.56
57.38
CrossDistill (Ours, 3-step)
83.15
58.05
Table 4 : Distillation-objective ablation at 97% attention sparsity (Wan2.1-T2V-14B-720P; 3-step CFG-free student; All 3-step student rows initialize from the same Stage-1 warm-up checkpoint.)
Figure 11 : Few-step distillation preserves sample diversity. On Wan2.1-T2V-14B-480P, for each of two prompts, we show one frame per video under three models (dense teacher; sparse model with 90% block sparsity and rank-64 compensation; 3-step student distilled from it) using the same four noise seeds.
Figure 12 : Wan2.1-T2V-1.3B-480P. Qualitative results at 480×832 resolution with 81 frames; TurboDiffusion and FastWan and SparkDiffusion operate at 90% attention sparsity.
Figure 13 : Wan2.1-T2V-14B-480P. Qualitative results at 480×832 resolution with 81 frames; all methods operate at 90% attention sparsity.
Figure 14 : Wan2.1-T2V-14B-720P. Qualitative results at 720×1280 resolution with 81 frames; TurboDiffusion and FastWan operate at 90% attention sparsity, SparkDiffusion at 97% .
Figure 15 : Wan2.1-I2V-14B-720P. Qualitative results at 720×1280 resolution with 81 frames; SparkDiffusion operates at 97% attention sparsity.
Figure 16 : Wan2.2-T2V-A14B-720P. Qualitative results at 720×1280 resolution with 81 frames; SparkDiffusion operates at 97% attention sparsity.
Figure 17 : Qualitative sample overlays for the same six toy sequence manifolds. In this toy setting, the multi-step sparse model can drift away from the data distribution, producing distorted or misaligned contours, while the 3-step distilled student produces more globally consistent shapes.
Model
Sp.
VBench (BF16)
VBench (FP8)
VBench-2.0 (BF16)
VBench-2.0 (FP8)
Wan2.1-T2V-1.3B-480P
90%
82.68
82.64
55.91
55.85
Wan2.1-T2V-14B-720P
90%
83.47
83.42
59.41
59.36
Wan2.1-T2V-14B-720P
97%
83.22
83.15
58.12
58.05
Wan2.2-T2V-A14B-720P
97%
83.41
83.36
58.53
58.46
Table 5: FP8 (W8A8) versus BF16 generation quality on deployed SparkDiffusion models (3-step CFG-free inference).
Figure 18 : Roadmap of the mechanism. Each box states one step of the argument and where it is proved. The proof is Hilbert-space geometry: a quadratic step-local loss and a linear terminal functional see the same residual differently. Sparsity enters only through empirical inputs—the residual left by sparse training and the trainable correction subspace—not as a mathematical object in the proof.
Symbol
Meaning
N,h,T
number of sampler steps, step size, horizon T=Nh
xk⋆
reference trajectory generated by v⋆
ek
per-step velocity error at step k
e=(e0,…,eN−1)
full velocity-error sequence
R(x)=c⊤x
linear terminal observable
Ψ
terminal error R(xN)−R(xN⋆)
Table 6: Notation used in the appendix. The key objects are the per-step velocity error e , the terminal sensitivity s , the trainable correction subspace U , and the two objectives: the step-local loss L and the terminal error Ψ .
Figure 19 : The exchange rate inside the surrogate (Corollary 7 , Proposition 8 ). Along the optimal path the terminal error falls linearly in the displacement fraction t while the step-local loss rises only quadratically ; the residual norm ∥e(a(t))∥h2=β2+2t2C follows the same quadratic. This is the formal reason for a staged recipe rather than a single joint objective. Both axes are surrogate quantities: neither is FID, sliced- W2 , or wall-clock training cost.
Sparse branch / recipe
Sp.
Val loss ↓
Term. error ↓
VBench ↑
VSA-style, step-local
90%
0.0892
0.1205
81.75
VSA-style, step-local
95%
0.0947
0.1232
80.89
VSA-style, step-local
97%
0.0998
0.1255
79.72
SLA-style, step-local
90%
0.0871
0.1203
81.87
SLA-style, step-local
95%
0.0932
0.1227
80.94
SLA-style, step-local
97%
0.0979
0.1251
79.84
Table 7: The high-sparsity trap across trainable sparse designs (Wan2.1-T2V-14B-480P). Step-local-only rows are trained from the same dense checkpoint with the flow-matching loss ( 2 ) under the extended 10,000-step budget of Sec. 3 ; terminal error is the paired latent MSE of Sec. 3 . “FastWan recipe” retrains the official VSA training pipeline [ 34 ] with only the sparsity level changed to 97% , and is evaluated under its official inference settings; “ + Stage 2” applies our trajectory-mixed distillation to the corresponding 97% checkpoint.
Hyperparameter
Setting
Model & Data
Backbone
Wan2.1-T2V-14B
Training dataset
OpenVid subset [ 14 ]
Number of training videos
2,000
Video setting
480×832 (480P), 720×1280 (720P), 81 frames
Sparse Architecture
Table 8: Training configuration of Stage 1 sparse warm-up on Wan2.1-T2V-14B.
Hyperparameter
Setting
Model setting
Student backbone
Wan2.1-T2V-14B
Student initialization
Stage 1 sparse warm-up checkpoint
Teacher model
Frozen dense Wan2.1-T2V-14B
Teacher CFG scale
5.0
Training dataset
Teacher-synthesized T2V dataset [ 38 ]
Table 9: Training configuration of Stage 2 trajectory-mixed distillation on Wan2.1-T2V-14B.