Organizations: Fudan University · Tencent HY · Shanghai Innovation Institute · MMLab, CUHK · Nankai University · Peking University · Shanghai Jiao Tong University
Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a 1.80× denoising speedup on Minimax-H3-Base and a 2.32× speedup on 3D asset generation, both with negligible quality loss.
Figures & tables
Figure 1: Quality and efficiency of MC-Sparse on Minimax-H3-Base at 768p. Left: quality metrics and DiT denoising latency. At 15% and 25% attention density, MC-Sparse achieves 1.80× and 1.61× speedup over dense attention (FA3), respectively. Right: frames generated at 15% density closely match the dense-attention reference across the video.
Figure 2: Oracle comparisons of the dense-to-sparse gap on Wan2.1-1.3B-T2V at 480p. Experimental settings are provided in Appendix A .
Figure 3: Three sources of the gap between dense and sparse attention.
Figure 4: Temporal stability and accuracy of reused exact selection.
Figure 5: Absolute temporal change of attention outputs and dense–sparse residuals under token-level selection. Lower values indicate smaller absolute variation across Δ denoising steps.
Figure 6: Pipeline overview of MC-Sparse.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
ImgQual ↑
BgCons ↑
Density ↓
Speedup ↑
Minimax-H3-Base, 768p, T2AV (video branch)
Full Attn (FA3)
-
-
-
69.20
91.72
100%
1.00×
Sol-Attn
23.66
0.816
0.234
68.83
91.73
32.4%
1.59×
PISA
23.45
0.811
0.243
68.54
91.64
30.0%
1.52×
MC-Sparse (density=25%)
28.44
0.899
0.161
69.00
91.81
25.0%
1.61×
MC-Sparse (density=15%)
27.30
0.882
0.176
68.94
91.85
15.0%
1.80×
Table 1: Video generation results. PSNR, SSIM, and LPIPS compare with dense-attention outputs; ImgQual and BgCons are VBench scores. Speedup measures DiT denoising only. Bold and underlined values indicate the best and second-best results, respectively, excluding Full Attn.
Figure 7: Qualitative comparison of video generation using Minimax-H3-Base.
Config
CD ↓
Vol-IoU-1536 ↑
F1@0.001 ↑
Uni3D-I ↑
ULIP3D-I ↑
Density ↓
Speedup ↑
HY3D-Internal
Full Attn (FA3)
-
-
-
0.3309
0.1204
100%
1.00×
Sol-Attn
1.901
50.55
70.36
0.3301
0.1203
36.1%
1.57×
PISA
0.976
63.22
83.29
0.3304
0.1204
25.0%
<1.87×
MC-Sparse (Ours)
0.177
82.91
96.33
0.3311
0.1206
15.0%
2.32×
Table 2: Image-to-geometry results on HY3D-Internal. Geometric fidelity is measured against dense-attention outputs; Uni3D-I and ULIP3D-I measure input-image consistency. Speedup measures DiT denoising only.
Figure 8: Qualitative comparison on 3D asset generation.
Configuration
Density s=0.2
Density s=0.25
Density s=0.3
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Vanilla BSA
20.96
0.7535
0.2714
21.67
0.7750
0.2469
22.26
0.7923
0.2278
+ exact KV selection (w. reuse)
21.60
0.7731
0.2482
22.63
0.8004
0.2176
23.47
0.8205
0.1956
+ token granularity
22.57
0.7954
0.2254
23.54
0.8188
0.1995
24.38
0.8380
0.1789
+ query grouping
23.51
0.8175
0.2017
24.47
0.8386
0.1784
25.24
0.8539
0.1624
+ compensation (full model)
27.05
0.8809
0.1326
27.35
0.8843
0.1294
27.54
0.8862
0.1273
Table 3: Cumulative ablation on Wan2.1-1.3B-T2V at 480p. Each row adds one component to the preceding configuration. Settings are provided in Appendix A .
Query grouping
KV granularity
Align.
Attention recall ↑
Output rel. L1↓
s=0.1
s=0.2
s=0.3
s=0.1
s=0.2
s=0.3
No reorder
block
1.00
0.637
0.778
0.847
0.176
0.101
0.069
token
1.00
0.745
0.857
0.910
0.130
0.070
0.044
k-means
block
2.08
0.743
0.830
0.874
0.160
0.098
0.071
token
1.45
0.840
0.904
0.935
0.093
0.054
0.036
block †
2.08
0.834
0.906
0.940
0.096
0.052
0.033
Table 4: KV-selection granularity and query grouping at matched realized attention density s . Align. is the ratio of executed to nominal sparse interactions after tile alignment. Black rows include this inflation in the density budget, whereas gray rows ignore it † .
s=0.20
s=0.25
s=0.30
Configuration
PSNR ↑
Δ
PSNR ↑
Δ
PSNR ↑
Δ
MC-Sparse w/o residual compensation
23.51
–
24.47
–
25.24
–
MC-Sparse
27.05
+3.54
27.35
+2.88
27.54
+2.30
Vanilla BSA
20.96
–
21.67
–
22.26
–
+ our residual compensation
22.21
+1.25
23.12
+1.45
23.90
+1.64
PISA
21.66
–
22.14
–
22.52
–
Table 5: PSNR (dB) before and after adding our residual compensation at target attention density s . Δ is the PSNR gain from adding the residual. PISA retains its native block-statistics compensation in both configurations.
Kernel
Sparse layout
s=0.2
s=0.3
s=0.5
BSA
regular blocks
4.96 × (0.99)
3.30 × (0.99)
1.97 × (0.99)
PISA kernel
regular blocks
4.11 × (0.82)
2.84 × (0.85)
1.73 × (0.87)
SVG2 kernel
variable-length blocks
4.02 × (0.80)
2.70 × (0.81)
1.62 × (0.81)
MC-Sparse
tile-aligned Q, gathered KV
4.79 × ( 0.96 )
3.25 × ( 0.98 )
1.97 × ( 0.99 )
Table 6: Attention-kernel speedup on Wan2.1-14B-T2V (720p), using a Hopper GPU in BF16 ( H=40 , D=128 , S=75600 ). Entries show speedup ρ=tFA3/t (efficiency η=ρs relative to the ideal 1/s speedup). Selection and grouping are excluded.
1.3B, 480p
14B, 720p
Grouping
Single
Amort.
Single
Amort.
Flash-KMeans init.
32.0
2.26
295.1
20.43
+ step update
1.50
13.38
Fast PDDP
11.3
0.57
66.6
3.33
Table 7: Query-grouping latency on Wan2.1-T2V (ms per layer). Amortized over 40 steps: 2 Fast PDDP calls or Flash-KMeans with a 50-iteration initialization and 39 two-iteration updates.
Model
V (MB)
C (TFLOP)
C/V (FLOP/B)
Transfer (ms)
Compute (ms)
Transfer hidden
Wan2.1-14B-T2V
2435
29.1
11,409
47.7
208.8
✓
HunyuanVideo-13B
2710
34.6
12,164
52.9
244.5
✓
Table 8: Per-layer cost of offloading cached output residuals and selected KV indices to CPU memory. V is the transfer volume and C the sparse-attention FLOPs; C/V gives the break-even compute-to-PCIe-bandwidth ratio for overlap. Check marks indicate transfers hidden by prefetching one layer ahead in the measured setting.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Setting
Warm-up layers
Warm-up steps
Anchor steps
Wan2.1-1.3B-T2V
480p, 5s, text-to-video
1/30
10/50
{10,26}
Wan2.1-14B-I2V
720p, 5s, image-to-video
1/40
10/50
{10,26}
Wan2.1-14B-T2V
720p, 5s, text-to-video
1/40
10/50
{10,26}
Hunyuan-13B
720p, 5.25s, text-to-video
1/20, 1/40
10/50
{10,26}
Minimax-H3-Base
768p, 14.4s, text-to-audio-video
1/52
10/49
{10,26}
HY3D-Internal
1536 resolution, image-to-geometry
1/40
1/12
{1}
Appendix
Table 9: Dense warm-up and MC-Sparse anchor configurations. Warm-up entries denote count / total count. Hunyuan lists dual-stream and single-stream layers, respectively.
TopP
Model
Qc
Kc
SVG2
SVG-EAR
Wan2.1-14B-T2V
300
1000
0.90
0.85
Wan2.1-14B-I2V
300
1000
0.90
0.85
Appendix
Table 10: SVG2 and SVG-EAR settings following the SVG-EAR evaluation setup. Qc and Kc denote the numbers of query and key clusters.
Figure 9: Qualitative comparison with full attention across video generation models.
Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs prohibitive computational cost due to the quadratic complexity of attention. Block sparse attention is a common approach to mitigate this by focusing computation on important regions. However, attention maps in DiTs exhibit inherently dynamic and fine-grained sparsity, which causes existing block sparse attention methods to degrade significantly in quality, especially at high sparsity ratios. In this paper, we revisit block sparse attention and derive a theoretical lower bound on attention recall to characterize the key factors governing its effectiveness. Guided by these insights, we propose DFSAttn, a training-free sparse attention framework that enables dynamic, fine-grained sparsification efficiently. DFSAttn incorporates three core designs: Hilbert curve-based token reordering to achieve fine-grained sparsity while preserving efficient GPU execution, hierarchical block scoring for accurate block importance estimation, and sparse mask caching with adaptive ratios to balance accuracy and efficiency. Experimental results demonstrate that DFSAttn consistently outperforms prior methods under high sparsity, achieving up to 2.1× end-to-end speedup while maintaining high generation quality. Our code is open-sourced and available at https://github.com/jessica-hujie/DFSAttn.
Sparse attention accelerates Diffusion Transformers (DiTs) for video generation by computing only the important tokens while skipping the rest. The token selection strategy is key to balancing sparsity and accuracy. We formulate the token filtering process as a dual-goal optimization problem: maximizing sparsity and minimizing accuracy degradation. Existing algorithms cannot fulfill both objectives simultaneously. For example, Top-p only considers the accuracy constraint, while Top-k maintains a fixed computational budget but loosens the accuracy constraint. This paper demonstrates that maintaining a fixed recall rate is sufficient for ensuring accuracy, whereas a fixed threshold is suboptimal for reducing computational cost. Therefore, we propose a dynamic thresholding scheme to improve sparsity while maintaining the same level of accuracy. Furthermore, our algorithm is deeply integrated with Flash Attention (FA), eliminating the need for any additional masking computation overhead. Experimental results on Wan 2.2 validate that, compared to the BLASST algorithm which is also integrated with FA, our dynamic thresholding strategy enhances sparsity from 61.42% to 82% with a VBench metric drop of less than 5%. This results in an approximate 15% in attention computation and a 1.61× increase in computational efficiency, which is 1.18x higher than that of BLASST.
While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing sparse methods face a dilemma: dynamic pruning suffers from prohibitive runtime overhead and memory fragmentation, while static heuristics fail to capture fine-grained dependencies. In this work, we propose ScalingAttention, a training-free framework grounded in a key inductive bias: while individual activations are input-dependent, the high-mass attention regions for each head rapidly converge to a stable, prompt-agnostic Intrinsic Sparse Topology. This topology is weight-encoded, scale-invariant, and efficient to extract. ScalingAttention decouples topology discovery from sparsity control via: (1) WEST (Weight-Encoded Sparse Topology), which extracts a robust block-sparse prior mask offline to eliminate runtime search; (2) FAST (Fidelity-Aware Sensitivity Tuning), which adaptively tunes head-wise sparsity based on diffusion fidelity requirements. To ensure practical acceleration, we co-design a hardware-aligned bit-wise block-sparse kernel. Experiments on Wan2.1 show up to 1.90X end-to-end speedup with superior fidelity, establishing a new Pareto frontier over state-of-the-art baselines.
Ruiliang Zhou, Xuecheng Wu, Kang He +6
KlingAI Research · Beijing Institute of Technology · NVIDIA +1