Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57× and 1.37× end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.
Figures & tables
Figure 1
Figure 3 : An overview of the PSA framework. (a) Unified Representation: The parametric model M captures diverse per-head sparse attention patterns, covering both intra-frame and inter-frame structures with varying widths, offsets, and strides. (b) Efficient Kernel: A dedicated CUDA kernel enables high-performance, hardware-aware computation across all supported sparse patterns. (c) Automatic Search: PSA-Search automatically identifies the optimal sparse configuration for each head by maximizing sparsity subject to a predefined output degradation threshold.
Figure 4 : Evaluation of the similarity of attention heatmaps across all 11 evaluation dimensions of VBench, where prompts 1–99 are compared against prompt 0 at the same head positions within a search space of 100(timesteps)×40(layers)×40(heads) .
Table 3 : Sensitivity tests on quality and kernel performance.
Method
Spar (%)
Speedup ↑
PSNR ↑
SVG
56.98
1.39 ×
23.00
STA
58.37
1.48 ×
12.52
SVG2
63.23
1.43 ×
25.80
Ours-min
56.47
1.51 ×
28.73
Ours-max
63.84
1.65 ×
26.76
Table 4 : Bracketing comparison on HunyuanVideo ( 768×1280 ). Ours-min and Ours-max are two operating points whose sparsity strictly brackets that of all baselines.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Mask Configuration
Head Count
(Tw,To,Ts)
(Sw,So,Ss)
# Heads
% Total
full window
32000
20.00
(7, 0, ∞ )
(3, 0, 3)
25257
15.79
(1, 0, ∞ )
(5, 0, 6)
11778
7.36
(7, 0, ∞ )
(1, 0, 3)
11069
6.92
(7, 0, ∞ )
(3, 0, 1)
9014
5.63
Appendix
Table 7 : Top 10 most frequent PSA configurations discovered by the search algorithm in Wan 2.1 14B using a search threshold of α=0.007 .
Mask Configuration
Head Count
(Tw,To,Ts)
(Sw,So,Ss)
# Heads
% Total
full window
14404
20.00
(1, 0, ∞ )
(10, 0, 6)
9213
12.80
(5, 0, ∞ )
(1, 0, 1)
5494
7.63
(5, 0, ∞ )
(1, 0, 3)
4889
6.79
(5, 0, ∞ )
(3, 0, 3)
4651
6.46
Appendix
Table 8 : Top 10 most frequent PSA configurations discovered for the HunyuanVideo model (768 × 1280 resolution) using a search threshold of α=0.018 .
Mask Configuration
Head Count
(Tw,To,Ts)
(Sw,So,Ss)
# Heads
% Total
full window
14400
20.00
(11, 0, ∞ )
(1, 0, 1)
6223
8.64
(1, 0, ∞ )
(5, 0, 6)
4937
6.86
(3, 2, ∞ )
(5, 0, 6)
3995
5.55
(11, 0, ∞ )
(3, 0, 1)
3964
5.51
Appendix
Table 9 : Top 10 most frequent PSA configurations for the HunyuanVideo model (720 × 1280 resolution) using a search threshold of α=0.03 .
Figure 5 : Empirical support for mλ∗ in RoPE-3D
Figure 6 : CogVideoX Heatmap Patterns.
Figure 7 : Heatmap Comparison Between Prompts.
Figure 8 : Sparsity variation with different α values across different models.
HunyuanVideo ( 720×1280 )
HunyuanVideo ( 768×1280 )
Wan 2.1 14B
Sparsity (%)
α
Sparsity (%)
α
Sparsity (%)
α
18.72
0.0003
20.59
0.0003
19.65
0.0003
28.65
0.0010
30.04
0.0010
31.63
0.0010
40.85
0.0030
42.16
0.0030
42.05
0.0021
49.73
0.0070
51.66
0.0070
51.51
0.0041
62.80
0.0600
60.33
0.0180
61.22
0.0110
Appendix
Table 10 : Alpha-sparsity correspondence for different models.
Method
Sparsity
E2E(s)
PSNR ↑
SSIM ↑
LPIPS ↓
SVG
10.48%
1393.25
30.34
0.9144
0.1247
30.39%
1116.87
27.68
0.8862
0.1552
50.32%
865.78
25.67
0.8494
0.1914
70.16%
622.37
21.21
0.7771
0.2651
SVG2
10.49%
1333.93
31.87
0.9076
0.1206
30.72%
1129.91
26.75
0.7753
0.2040
Appendix
Table 11 : End-to-end performance and quality comparison across varying sparsity on HunyuanVideo (768 × 1280).
Figure 9 : HunyuanVideo(720 × 1280) FA3 Baseline vs. Ours.
Figure 10 : HunyuanVideo(768 × 1280) FA3 Baseline vs. Ours.
Figure 11 : Wan 2.1(720 × 1280) FA3 Baseline vs. Ours.