Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57× and 1.37× end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.
Figures & tables
Figure 1
Figure 3 : An overview of the PSA framework. (a) Unified Representation: The parametric model M captures diverse per-head sparse attention patterns, covering both intra-frame and inter-frame structures with varying widths, offsets, and strides. (b) Efficient Kernel: A dedicated CUDA kernel enables high-performance, hardware-aware computation across all supported sparse patterns. (c) Automatic Search: PSA-Search automatically identifies the optimal sparse configuration for each head by maximizing sparsity subject to a predefined output degradation threshold.
Figure 4 : Evaluation of the similarity of attention heatmaps across all 11 evaluation dimensions of VBench, where prompts 1–99 are compared against prompt 0 at the same head positions within a search space of 100(timesteps)×40(layers)×40(heads) .
Table 3 : Sensitivity tests on quality and kernel performance.
Method
Spar (%)
Speedup ↑
PSNR ↑
SVG
56.98
1.39 ×
23.00
STA
58.37
1.48 ×
12.52
SVG2
63.23
1.43 ×
25.80
Ours-min
56.47
1.51 ×
28.73
Ours-max
63.84
1.65 ×
26.76
Table 4 : Bracketing comparison on HunyuanVideo ( 768×1280 ). Ours-min and Ours-max are two operating points whose sparsity strictly brackets that of all baselines.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Mask Configuration
Head Count
(Tw,To,Ts)
(Sw,So,Ss)
# Heads
% Total
full window
32000
20.00
(7, 0, ∞ )
(3, 0, 3)
25257
15.79
(1, 0, ∞ )
(5, 0, 6)
11778
7.36
(7, 0, ∞ )
(1, 0, 3)
11069
6.92
(7, 0, ∞ )
(3, 0, 1)
9014
5.63
Appendix
Table 7 : Top 10 most frequent PSA configurations discovered by the search algorithm in Wan 2.1 14B using a search threshold of α=0.007 .
Mask Configuration
Head Count
(Tw,To,Ts)
(Sw,So,Ss)
# Heads
% Total
full window
14404
20.00
(1, 0, ∞ )
(10, 0, 6)
9213
12.80
(5, 0, ∞ )
(1, 0, 1)
5494
7.63
(5, 0, ∞ )
(1, 0, 3)
4889
6.79
(5, 0, ∞ )
(3, 0, 3)
4651
6.46
Appendix
Table 8 : Top 10 most frequent PSA configurations discovered for the HunyuanVideo model (768 × 1280 resolution) using a search threshold of α=0.018 .
Mask Configuration
Head Count
(Tw,To,Ts)
(Sw,So,Ss)
# Heads
% Total
full window
14400
20.00
(11, 0, ∞ )
(1, 0, 1)
6223
8.64
(1, 0, ∞ )
(5, 0, 6)
4937
6.86
(3, 2, ∞ )
(5, 0, 6)
3995
5.55
(11, 0, ∞ )
(3, 0, 1)
3964
5.51
Appendix
Table 9 : Top 10 most frequent PSA configurations for the HunyuanVideo model (720 × 1280 resolution) using a search threshold of α=0.03 .
Figure 5 : Empirical support for mλ∗ in RoPE-3D
Figure 6 : CogVideoX Heatmap Patterns.
Figure 7 : Heatmap Comparison Between Prompts.
Figure 8 : Sparsity variation with different α values across different models.
HunyuanVideo ( 720×1280 )
HunyuanVideo ( 768×1280 )
Wan 2.1 14B
Sparsity (%)
α
Sparsity (%)
α
Sparsity (%)
α
18.72
0.0003
20.59
0.0003
19.65
0.0003
28.65
0.0010
30.04
0.0010
31.63
0.0010
40.85
0.0030
42.16
0.0030
42.05
0.0021
49.73
0.0070
51.66
0.0070
51.51
0.0041
62.80
0.0600
60.33
0.0180
61.22
0.0110
Appendix
Table 10 : Alpha-sparsity correspondence for different models.
Method
Sparsity
E2E(s)
PSNR ↑
SSIM ↑
LPIPS ↓
SVG
10.48%
1393.25
30.34
0.9144
0.1247
30.39%
1116.87
27.68
0.8862
0.1552
50.32%
865.78
25.67
0.8494
0.1914
70.16%
622.37
21.21
0.7771
0.2651
SVG2
10.49%
1333.93
31.87
0.9076
0.1206
30.72%
1129.91
26.75
0.7753
0.2040
Appendix
Table 11 : End-to-end performance and quality comparison across varying sparsity on HunyuanVideo (768 × 1280).
Figure 9 : HunyuanVideo(720 × 1280) FA3 Baseline vs. Ours.
Figure 10 : HunyuanVideo(768 × 1280) FA3 Baseline vs. Ours.
Figure 11 : Wan 2.1(720 × 1280) FA3 Baseline vs. Ours.
Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs prohibitive computational cost due to the quadratic complexity of attention. Block sparse attention is a common approach to mitigate this by focusing computation on important regions. However, attention maps in DiTs exhibit inherently dynamic and fine-grained sparsity, which causes existing block sparse attention methods to degrade significantly in quality, especially at high sparsity ratios. In this paper, we revisit block sparse attention and derive a theoretical lower bound on attention recall to characterize the key factors governing its effectiveness. Guided by these insights, we propose DFSAttn, a training-free sparse attention framework that enables dynamic, fine-grained sparsification efficiently. DFSAttn incorporates three core designs: Hilbert curve-based token reordering to achieve fine-grained sparsity while preserving efficient GPU execution, hierarchical block scoring for accurate block importance estimation, and sparse mask caching with adaptive ratios to balance accuracy and efficiency. Experimental results demonstrate that DFSAttn consistently outperforms prior methods under high sparsity, achieving up to 2.1× end-to-end speedup while maintaining high generation quality. Our code is open-sourced and available at https://github.com/jessica-hujie/DFSAttn.
While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing sparse methods face a dilemma: dynamic pruning suffers from prohibitive runtime overhead and memory fragmentation, while static heuristics fail to capture fine-grained dependencies. In this work, we propose ScalingAttention, a training-free framework grounded in a key inductive bias: while individual activations are input-dependent, the high-mass attention regions for each head rapidly converge to a stable, prompt-agnostic Intrinsic Sparse Topology. This topology is weight-encoded, scale-invariant, and efficient to extract. ScalingAttention decouples topology discovery from sparsity control via: (1) WEST (Weight-Encoded Sparse Topology), which extracts a robust block-sparse prior mask offline to eliminate runtime search; (2) FAST (Fidelity-Aware Sensitivity Tuning), which adaptively tunes head-wise sparsity based on diffusion fidelity requirements. To ensure practical acceleration, we co-design a hardware-aligned bit-wise block-sparse kernel. Experiments on Wan2.1 show up to 1.90X end-to-end speedup with superior fidelity, establishing a new Pareto frontier over state-of-the-art baselines.
Ruiliang Zhou, Xuecheng Wu, Kang He +6
KlingAI Research · Beijing Institute of Technology · NVIDIA +1
Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse attention methods degrade under high sparsity. We show empirically that generation quality is determined not by the sparsity ratio itself, but by how well the sparse mask aligns with the tile-wise geometry of full attention. Based on this insight, we propose Veda, a distilled sparse attention framework that formulates tile selection as an explicit reconstruction problem from full attention. Veda integrates statistics-aware tile scoring with head-aware tiling to reduce estimation error and structural mismatch, enabling aggressive sparsity. A hardware-efficient tile-skipping kernel converts theoretical sparsity into practical wall-clock speedups. Experiments on large video diffusion models, including Waver and Wan2.1, demonstrate substantial acceleration with no noticeable degradation in generation quality. To generate 720P 10-second videos on Waver-T2V-12B, Veda achieves a 5.1× end-to-end speedup and a 10.5× self-attention speedup, reducing attention overhead from 92% to 50%. Notably, the gains increase with sequence length, indicating that Veda scales favorably with spatiotemporal resolution across models.
Shihao Han, Hao Yang, Xinting Hu +3
ByteDance Inc. · University of Science and Technology of China · The University of Hong Kong