Organizations: Fudan University · Tencent HY · Shanghai Innovation Institute · MMLab, CUHK · Nankai University · Peking University · Shanghai Jiao Tong University
Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a 1.80× denoising speedup on Minimax-H3-Base and a 2.32× speedup on 3D asset generation, both with negligible quality loss.
Figures & tables
Figure 1: Quality and efficiency of MC-Sparse on Minimax-H3-Base at 768p. Left: quality metrics and DiT denoising latency. At 15% and 25% attention density, MC-Sparse achieves 1.80× and 1.61× speedup over dense attention (FA3), respectively. Right: frames generated at 15% density closely match the dense-attention reference across the video.
Figure 2: Oracle comparisons of the dense-to-sparse gap on Wan2.1-1.3B-T2V at 480p. Experimental settings are provided in Appendix A .
Figure 3: Three sources of the gap between dense and sparse attention.
Figure 4: Temporal stability and accuracy of reused exact selection.
Figure 5: Absolute temporal change of attention outputs and dense–sparse residuals under token-level selection. Lower values indicate smaller absolute variation across Δ denoising steps.
Figure 6: Pipeline overview of MC-Sparse.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
ImgQual ↑
BgCons ↑
Density ↓
Speedup ↑
Minimax-H3-Base, 768p, T2AV (video branch)
Full Attn (FA3)
-
-
-
69.20
91.72
100%
1.00×
Sol-Attn
23.66
0.816
0.234
68.83
91.73
32.4%
1.59×
PISA
23.45
0.811
0.243
68.54
91.64
30.0%
1.52×
MC-Sparse (density=25%)
28.44
0.899
0.161
69.00
91.81
25.0%
1.61×
MC-Sparse (density=15%)
27.30
0.882
0.176
68.94
91.85
15.0%
1.80×
Table 1: Video generation results. PSNR, SSIM, and LPIPS compare with dense-attention outputs; ImgQual and BgCons are VBench scores. Speedup measures DiT denoising only. Bold and underlined values indicate the best and second-best results, respectively, excluding Full Attn.
Figure 7: Qualitative comparison of video generation using Minimax-H3-Base.
Config
CD ↓
Vol-IoU-1536 ↑
F1@0.001 ↑
Uni3D-I ↑
ULIP3D-I ↑
Density ↓
Speedup ↑
HY3D-Internal
Full Attn (FA3)
-
-
-
0.3309
0.1204
100%
1.00×
Sol-Attn
1.901
50.55
70.36
0.3301
0.1203
36.1%
1.57×
PISA
0.976
63.22
83.29
0.3304
0.1204
25.0%
<1.87×
MC-Sparse (Ours)
0.177
82.91
96.33
0.3311
0.1206
15.0%
2.32×
Table 2: Image-to-geometry results on HY3D-Internal. Geometric fidelity is measured against dense-attention outputs; Uni3D-I and ULIP3D-I measure input-image consistency. Speedup measures DiT denoising only.
Figure 8: Qualitative comparison on 3D asset generation.
Configuration
Density s=0.2
Density s=0.25
Density s=0.3
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Vanilla BSA
20.96
0.7535
0.2714
21.67
0.7750
0.2469
22.26
0.7923
0.2278
+ exact KV selection (w. reuse)
21.60
0.7731
0.2482
22.63
0.8004
0.2176
23.47
0.8205
0.1956
+ token granularity
22.57
0.7954
0.2254
23.54
0.8188
0.1995
24.38
0.8380
0.1789
+ query grouping
23.51
0.8175
0.2017
24.47
0.8386
0.1784
25.24
0.8539
0.1624
+ compensation (full model)
27.05
0.8809
0.1326
27.35
0.8843
0.1294
27.54
0.8862
0.1273
Table 3: Cumulative ablation on Wan2.1-1.3B-T2V at 480p. Each row adds one component to the preceding configuration. Settings are provided in Appendix A .
Query grouping
KV granularity
Align.
Attention recall ↑
Output rel. L1↓
s=0.1
s=0.2
s=0.3
s=0.1
s=0.2
s=0.3
No reorder
block
1.00
0.637
0.778
0.847
0.176
0.101
0.069
token
1.00
0.745
0.857
0.910
0.130
0.070
0.044
k-means
block
2.08
0.743
0.830
0.874
0.160
0.098
0.071
token
1.45
0.840
0.904
0.935
0.093
0.054
0.036
block †
2.08
0.834
0.906
0.940
0.096
0.052
0.033
Table 4: KV-selection granularity and query grouping at matched realized attention density s . Align. is the ratio of executed to nominal sparse interactions after tile alignment. Black rows include this inflation in the density budget, whereas gray rows ignore it † .
s=0.20
s=0.25
s=0.30
Configuration
PSNR ↑
Δ
PSNR ↑
Δ
PSNR ↑
Δ
MC-Sparse w/o residual compensation
23.51
–
24.47
–
25.24
–
MC-Sparse
27.05
+3.54
27.35
+2.88
27.54
+2.30
Vanilla BSA
20.96
–
21.67
–
22.26
–
+ our residual compensation
22.21
+1.25
23.12
+1.45
23.90
+1.64
PISA
21.66
–
22.14
–
22.52
–
Table 5: PSNR (dB) before and after adding our residual compensation at target attention density s . Δ is the PSNR gain from adding the residual. PISA retains its native block-statistics compensation in both configurations.
Kernel
Sparse layout
s=0.2
s=0.3
s=0.5
BSA
regular blocks
4.96 × (0.99)
3.30 × (0.99)
1.97 × (0.99)
PISA kernel
regular blocks
4.11 × (0.82)
2.84 × (0.85)
1.73 × (0.87)
SVG2 kernel
variable-length blocks
4.02 × (0.80)
2.70 × (0.81)
1.62 × (0.81)
MC-Sparse
tile-aligned Q, gathered KV
4.79 × ( 0.96 )
3.25 × ( 0.98 )
1.97 × ( 0.99 )
Table 6: Attention-kernel speedup on Wan2.1-14B-T2V (720p), using a Hopper GPU in BF16 ( H=40 , D=128 , S=75600 ). Entries show speedup ρ=tFA3/t (efficiency η=ρs relative to the ideal 1/s speedup). Selection and grouping are excluded.
1.3B, 480p
14B, 720p
Grouping
Single
Amort.
Single
Amort.
Flash-KMeans init.
32.0
2.26
295.1
20.43
+ step update
1.50
13.38
Fast PDDP
11.3
0.57
66.6
3.33
Table 7: Query-grouping latency on Wan2.1-T2V (ms per layer). Amortized over 40 steps: 2 Fast PDDP calls or Flash-KMeans with a 50-iteration initialization and 39 two-iteration updates.
Model
V (MB)
C (TFLOP)
C/V (FLOP/B)
Transfer (ms)
Compute (ms)
Transfer hidden
Wan2.1-14B-T2V
2435
29.1
11,409
47.7
208.8
✓
HunyuanVideo-13B
2710
34.6
12,164
52.9
244.5
✓
Table 8: Per-layer cost of offloading cached output residuals and selected KV indices to CPU memory. V is the transfer volume and C the sparse-attention FLOPs; C/V gives the break-even compute-to-PCIe-bandwidth ratio for overlap. Check marks indicate transfers hidden by prefetching one layer ahead in the measured setting.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Setting
Warm-up layers
Warm-up steps
Anchor steps
Wan2.1-1.3B-T2V
480p, 5s, text-to-video
1/30
10/50
{10,26}
Wan2.1-14B-I2V
720p, 5s, image-to-video
1/40
10/50
{10,26}
Wan2.1-14B-T2V
720p, 5s, text-to-video
1/40
10/50
{10,26}
Hunyuan-13B
720p, 5.25s, text-to-video
1/20, 1/40
10/50
{10,26}
Minimax-H3-Base
768p, 14.4s, text-to-audio-video
1/52
10/49
{10,26}
HY3D-Internal
1536 resolution, image-to-geometry
1/40
1/12
{1}
Appendix
Table 9: Dense warm-up and MC-Sparse anchor configurations. Warm-up entries denote count / total count. Hunyuan lists dual-stream and single-stream layers, respectively.
TopP
Model
Qc
Kc
SVG2
SVG-EAR
Wan2.1-14B-T2V
300
1000
0.90
0.85
Wan2.1-14B-I2V
300
1000
0.90
0.85
Appendix
Table 10: SVG2 and SVG-EAR settings following the SVG-EAR evaluation setup. Qc and Kc denote the numbers of query and key clusters.
Figure 9: Qualitative comparison with full attention across video generation models.