Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval can either degrade generation quality or yield unnecessary computation. We identify two retrieval mismatches in methods that retrieve blocks using the averaged representations of query and key blocks: (i) query-side aggregation mismatch, where averaging queries before Softmax fails to preserve their individual attention preferences, and (ii) key-side clustering metric mismatch, where standard Euclidean clustering in the original key space can group keys with dissimilar QK scores under the current query, so their average representation may not accurately represent how the current query scores individual keys. These mismatches can lead to inaccurate block retrieval. To address these mismatches, we propose PARK, a training-free sparse attention method for accurate block retrieval. PARK retains every original query, independently normalizes its attention over key blocks, and then averages these distributions within each query block. It also uses information from the current queries to transform keys before clustering, so that keys receiving similar QK scores are grouped together. A fused GPU kernel further reduces the overhead of block retrieval. Experiments on HunyuanVideo and Wan demonstrate that PARK improves block retrieval accuracy and preserves generation quality while accelerating inference, achieving the best quality-efficiency trade-off among the compared sparse attention methods.
Figures & tables
Figure 1: Two-sided retrieval mismatches and approximation errors in centroid-based block retrieval. (a) Averaging queries into a single centroid before Softmax normalization ignores their individual key preferences, leading to inaccurate block retrieval. (b) Standard Euclidean clustering considers only distances between keys in feature space and does not ensure that keys within the same cluster receive similar scores from the current query, leading to inaccurate key-centroid estimates. (c) Retaining the original queries instead of using query centroids (top) and clustering keys by QK-score similarity (bottom) both reduce approximation errors. More experimental details are provided in Appendix A .
Ordering
Query
Key
HunyuanVideo ( Kong et al. (2024) )
Wan 2.1-14B ( Wan et al. (2025) )
Mean
P05
Worst
Mean
P05
Worst
Full
Full
91.29
90.02
90.01
91.71
90.02
90.01
Centroid
Centroid
88.06
81.53
64.21
90.43
87.11
73.89
Centroid
Full
89.10
84.44
68.40
91.00
88.64
79.06
Clustered
Full
Centroid
90.70
87.62
81.64
91.56
89.36
83.61
Full
Full
90.64
90.02
90.01
90.89
90.03
90.02
Table 1: Top- p block-retrieval recall ( p=0.9 ) under different query/key representations. Mean, P05, and Worst denote the mean, fifth percentile, and minimum attention recall across all attention heads and layers, respectively. All values are percentages. Higher is better.
Figure 2: Overview of PARK. QRKC transforms keys using a metric derived from the current queries, so that keys receiving similar QK scores are grouped together, while queries are clustered using standard Euclidean K-means. PAMA retains every original query, computes its Softmax-normalized weights over the key centroids, and averages them within each query block to produce the block-importance map.
Model
Method
Quality
Efficiency
PSNR ↑
SSIM ↑
LPIPS ↓
SubCons. ↑
BackCons. ↑
ImageQual. ↑
AesQual. ↑
Latency ↓
Speedup ↑
Wan2.2-14B
Dense
–
–
–
97.13%
97.22%
73.78%
65.56%
3792s
1.00 ×
SpargeAttn
18.32
0.714
0.295
97.10%
97.33%
73.73%
65.55%
2614s
1.45 ×
XAttention
17.31
0.640
0.353
96.89%
97.10%
73.87%
66.07%
2819s
1.34 ×
SVG
17.33
0.663
0.342
96.99%
96.96%
73.73%
65.85%
2668s
1.42 ×
SVG2
18.77
0.730
0.288
96.88%
97.11%
73.71%
65.20%
2672s
1.42 ×
Table 2: Quality and efficiency benchmark results for PARK and the baselines.
Model
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
Latency ↓
PARK
29.85
0.931
0.168
1827s
w/o PAMA
28.40
0.911
0.205
1824s
HunyuanVideo
w/o QRKC
29.54
0.928
0.173
1826s
PARK
23.90
0.871
0.189
2359s
w/o PAMA
23.48
0.864
0.199
2353s
Wan2.1-14B
w/o QRKC
23.80
0.869
0.190
2354s
Table 3: Ablation study of PAMA and QRKC on HunyuanVideo and Wan2.1-14B.
Figure 3: Examples of videos generated by PARK and SVG2 on HunyuanVideo and Wan2.1-14B.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Direct evaluation of the two retrieval mismatches. Left: retaining original queries reduces attention-weight approximation error compared with query-centroid estimation. Right: QRKC reduces centered log-mass error compared with standard Euclidean clustering. Results correspond to Figure 1 (c). Lower is better.
Query
Key
Wan2.1-14B
HunyuanVideo
Full
Full
0.0000
0.0000
Centroid
Centroid
0.1774
0.2092
Centroid
Full
0.1335
0.1667
Full
Centroid
0.0927
0.1066
Appendix
Table 4: Mean attention-weight TV under different query and key representations, using the same Euclidean partitions. Full-query, full-key estimation is the exact reference and has zero error by definition. The two key-centroid rows reproduce the M1 results in Figure 4 . Lower is better. Bold indicates the lowest error among the compressed representations.
Model
Metric
Euclidean
QRKC
Change
Wan2.1-14B
Raw-key MSE
66.88
72.59
+8.5%
QK-score squared error
11.49
7.28
−36.6%
HunyuanVideo
Raw-key MSE
79.51
91.31
+14.8%
QK-score squared error
2.92
2.26
−22.8%
Appendix
Table 5: Key-space and QK-score errors. Changes are relative to Euclidean clustering.
Figure 5: Effect of QRKC on mean and P05 attention recall.
Model
Implementation
Runtime
Additional GPU Memory
Latency ↓
Speedup ↑
Peak ↓
Reduction ↑
Saving ↑
HunyuanVideo
PyTorch
87.95ms
1.00 ×
21.09 GB
1.00 ×
0.00%
Triton (Two-pass)
28.12ms
3.13 ×
24.12 MB
895.60 ×
99.89%
Triton (Full- K )
15.82ms
5.56 ×
24.17 MB
893.86 ×
99.89%
TileLang (Two-pass)
21.35ms
4.12 ×
24.12 MB
895.60 ×
99.89%
TileLang (Full- K )
21.08ms
4.17 ×
24.17 MB
893.86 ×
99.89%
Appendix
Table 6: PAMA runtime and additional GPU memory. Speedup, memory reduction (multiplicative factor), and memory saving are relative to PyTorch.
Figure 6: Time breakdown of QRKC overhead excluding K-means clustering. The total measured overhead remains below 2 ms. We use the same experimental settings as in Table 6 .
NQ
NK
PSNR ↑
SSIM ↑
LPIPS ↓
Latency ↓
Speedup ↑
64
1,024
29.38
0.926
0.175
1817s
1.68 ×
128
1,024
29.85
0.931
0.168
1827s
1.67 ×
192
1,024
29.94
0.933
0.167
1841s
1.66 ×
256
1,024
30.13
0.934
0.164
1855s
1.64 ×
320
1,024
30.13
0.935
0.162
1875s
1.63 ×
128
512
29.56
0.928
0.174
1830s
1.67 ×
Appendix
Table 7: Sensitivity to the numbers of query and key centroids on HunyuanVideo.
NQ
NK
PSNR ↑
SSIM ↑
LPIPS ↓
Latency ↓
Speedup ↑
64
512
27.30
0.903
0.183
410s
1.82 ×
128
512
27.52
0.908
0.177
423s
1.76 ×
192
512
27.66
0.911
0.174
427s
1.74 ×
256
512
27.71
0.911
0.173
430s
1.73 ×
320
512
27.79
0.913
0.172
440s
1.69 ×
128
768
27.67
0.910
0.175
416s
1.79 ×
Appendix
Table 8: Sensitivity to the numbers of query and key centroids on Wan2.1-1.3B.
T
PSNR ↑
SSIM ↑
LPIPS ↓
Latency ↓
Speedup ↑
{10}
29.55
0.928
0.174
1829s
1.67 ×
{10,30}
29.67
0.930
0.171
1819s
1.68 ×
{10,20,30}
29.84
0.932
0.168
1819s
1.68 ×
{10,20,30,40}
29.85
0.931
0.168
1827s
1.67 ×
Appendix
Table 9: Effect of clustering and sparse-mask recomputation schedules on HunyuanVideo with 50 denoising steps. Speedup is measured relative to dense attention.
Figure 7: Efficiency-quality trade-off of PARK.
Ordering
Query
Key
HunyuanVideo
Wan 2.1-14B
Mean
P05
Worst
Mean
P05
Worst
Full
Full
77.05
27.41
18.97
73.48
32.81
19.66
Centroid
Centroid
75.37
27.27
18.96
72.35
32.78
19.66
Centroid
Full
76.14
27.38
18.96
72.88
32.80
19.66
Clustered
Full
Centroid
76.50
27.29
18.96
73.12
32.79
19.66
Full
Full
61.90
18.93
12.73
56.49
21.17
10.51
Appendix
Table 10: Top-k block-retrieval attention recall at a fixed retrieval ratio ( ρ=0.1 ) under different query/key representations. Mean, P05, and Worst denote the mean, fifth percentile, and minimum attention recall across all attention heads and layers, respectively. All values are percentages. Higher is better.
Figure 8: Comparison of Dense Attention and PARK for text-to-video generation with HunyuanVideo.
Figure 9: Comparison of Dense Attention and PARK for text-to-video generation with Wan2.1-14B.
Figure 10: Comparison of Dense Attention and PARK for text-to-video generation with Wan2.1-1.3B.
Figure 11: Comparison of Dense Attention and PARK for text-to-video generation with Wan2.2-14B.
Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs prohibitive computational cost due to the quadratic complexity of attention. Block sparse attention is a common approach to mitigate this by focusing computation on important regions. However, attention maps in DiTs exhibit inherently dynamic and fine-grained sparsity, which causes existing block sparse attention methods to degrade significantly in quality, especially at high sparsity ratios. In this paper, we revisit block sparse attention and derive a theoretical lower bound on attention recall to characterize the key factors governing its effectiveness. Guided by these insights, we propose DFSAttn, a training-free sparse attention framework that enables dynamic, fine-grained sparsification efficiently. DFSAttn incorporates three core designs: Hilbert curve-based token reordering to achieve fine-grained sparsity while preserving efficient GPU execution, hierarchical block scoring for accurate block importance estimation, and sparse mask caching with adaptive ratios to balance accuracy and efficiency. Experimental results demonstrate that DFSAttn consistently outperforms prior methods under high sparsity, achieving up to 2.1× end-to-end speedup while maintaining high generation quality. Our code is open-sourced and available at https://github.com/jessica-hujie/DFSAttn.
Video diffusion transformers (DiTs) suffer from prohibitive inference latency due to quadratic attention complexity. Existing sparse attention methods either overlook semantic similarity or fail to adapt to heterogeneous token distributions across layers, leading to model performance degradation. We propose AdaCluster, a training-free adaptive clustering framework that accelerates the generation of DiTs while preserving accuracy. AdaCluster applies an angle-similarity-preserving clustering method to query vectors for higher compression, and designs a euclidean-similarity-preserving clustering method for keys, covering cluster number assignment, threshold-wise adaptive clustering, and efficient critical cluster selection. Experiments on CogVideoX-2B, HunyuanVideo, and Wan-2.1 on one A40 GPU demonstrate up to 1.67-4.31x speedup with negligible quality degradation.
Haoyue Tan, Shengnan Wang, Yulin Qiao +5
University of Science and Technology of China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center · 3Independent Researcher +2
Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end-to-end gains. We present SPADE, a training-free sparse-attention engine of three parts: (i) vDiT-SSR, a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions; (ii) runtime scheme generation using SICS and a head-wise policy; and (iii) an executor with low-overhead index search, flash block-sparse attention, and kernel grouping. Across Hunyuan-Video and Wan 2.1/2.2 for text-to-video and image-to-video generation, SPADE raises sparsity and speed while preserving quality, accelerating attention by 2.26x-3.40x and end-to-end inference by 1.49x-1.80x. Our code is open-sourced at https://github.com/6somehow/DAC-SPADE.