Organizations: School of Artificial Intelligence, Shanghai Jiao Tong University · School of Data Science, Fudan University · School of Computing and Data Science, The University of Hong Kong
Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity. However, isolated windows block cross-window interaction, often introducing visible grid-like artifacts in the generated results. Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality. However, its irregular computation patterns create a substantial gap between theoretical and practical speedups and require specialized kernels tailored to each hardware backend. To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration. Specifically, BASA replaces visual self-attention with shifted local-window attention. By introducing a structured window-shifting scheme across DiT blocks, we allow tokens divided by window boundaries in one layer to communicate in the following layers, thereby achieving global information exchange and eliminating window-induced visual artifacts. Notably, our design introduces no additional irregular operators or customized kernels, making it readily deployable on existing attention backends and closing the gap between theoretical sparsity and practical acceleration. Experiments demonstrate that BASA achieves measured speedups exceeding 90% of the theoretical estimates on FLUX and delivers a 4.52× attention speedup on Wan while maintaining competitive generation quality. Codes are publicly available at: https://github.com/lama0110/BASA.
Figures & tables
Figure 1: Ultra-resolution results generated by FLUX.1-dev and Wan2.1-1.3B with our approach. Resolution is marked on the bottom-right corner of each result in the format of width×height. Corresponding prompts can be found in the appendix.
Figure 2: Motivation for BASA. Swin-style window attention ( Liu et al., 2021 ) provides fast inference but introduces severe window boundary artifacts. CLEAR ( Liu et al., 2025 ) preserves high visual fidelity, while its practical speed gain is relatively limited. BASA preserves faithful reconstruction while achieving efficient inference, providing a favorable quality–efficiency trade-off.
Figure 3: Overview of BASA. (a) The proposed sparse attention framework integrates multiple components to recover both local and global visual interactions under reduced attention cost. (b) BASA is applied as a drop-in replacement for the visual self-attention module in DiT blocks, while the remaining generation pipeline remains unchanged.
Figure 4: Qualitative comparison of image resolution extrapolation from 1024×1024 to 2048×2048 on FLUX.1-dev. Competing sparse-attention methods may exhibit window-boundary artifacts or duplicated structures in some cases, while BASA maintains or adds more consistent spatial details.
Method
Config
PSNR ↑
SSIM ↑
FID ↓
LPIPS ↓
CLIP-I ↑
DINO ↑
CLIP-T ↑
IS ↑
FLUX
–
–
–
–
–
–
–
31.13
23.41
CLEAR
r=16 , distilled
23.50
0.83
23.28
0.23
97.09
94.94
31.16
23.44
Swin
w=32 , distilled
23.33
0.83
34.10
0.28
93.62
91.62
31.43
23.67
STA
w=32 , distilled
23.51
0.83
24.02
0.24
96.45
94.57
31.29
23.68
BASA
w=8,p=4 , distilled
21.32
0.78
38.65
0.31
92.61
87.59
31.48
22.50
BASA
w=16,p=4 , distilled
22.57
0.82
26.67
0.25
96.45
93.53
31.34
24.10
Table 1: Quantitative comparison for image resolution extrapolation from 1024×1024 to 2048×2048 with FLUX.1-dev. PSNR, SSIM, FID, LPIPS, CLIP-I, and DINO are computed against the corresponding dense-FLUX output.
Method
Config.
Implementation
Latency (ms) ↓
Act. ↑
Theo. ↑
Real. ↑
FLUX
Dense
SDPA
21.666
1.000
1.000
1.000
Swin
w=32
SDPA
8.819
2.455
3.187
0.770
CLEAR
r=16
FlexAttn
13.106
1.629
3.229
0.504
STA
w=32
Triton
12.955
1.648
3.092
0.533
BASA
w=8,p=4
SDPA
8.079
2.680
3.064
0.874
BASA
w=16,p=4
SDPA
8.131
2.665
2.987
0.892
Table 2: Theoretical and measured weighted-MSA efficiency for 1024×1024 to 2048×2048 resolution extrapolation. Weighted MSA denotes the average full multi-head self-attention latency per layer, weighted over all double- and single-stream attention layers of FLUX and the corresponding window schedules. Backend indicates the attention execution backend used in our implementation. Act., Theo., and Real. denote measured speedup, theoretical speedup, and their ratio, respectively.
Method
Config (latent)
Impl.
Sparsity
TFLOPs (Attn.)
Latency (ms)
MFU
Kernel Eff.
Speedup
Wan
–
SDPA
0.00%
106.12
509.66
66.74%
100.00%
1.00×
CLEAR
w=16
Flex
89.46%
11.18
485.79
7.38%
11.05%
1.05×
Swin
w=16
Flex
96.25%
3.84
396.60
3.10%
4.65%
1.29×
STA
w=(21,30,39),t=(3,10,13)
Triton
81.25%
19.90
175.09
36.42%
53.48%
2.85×
BASA
w=(10,10),p=(2,2)
SDPA
70.17%
31.66
255.87
39.66%
59.42%
1.99×
BASA
w=(15,26),p=(4,4)
SDPA
81.25%
19.90
126.07
50.59%
75.80%
4.04×
Table 3: Attention efficiency for video resolution extrapolation from 832×480 to 1664×960 with Wan2.1-T2V-1.3B. Configurations are specified on the latent grid. Latency measures attention-only forward time after warm-up, and speedup is reported relative to dense Wan.
Method
Config
SC ↑
BC ↑
MS ↑
DD ↑
AQ ↑
IQ ↑
Wan
–
96.43%
96.84%
99.18%
0.37
57.53%
65.22%
CLEAR
w=(8,8),d=4
96.27%
96.54%
99.19%
0.35
57.50%
67.11%
Swin
w=(8,8)
93.62%
94.59%
98.15%
0.04
41.36%
60.90%
STA
w=(18,24,24),t=(6,8,8)
95.89%
96.05%
99.16%
0.37
56.26%
67.56%
BASA
w=(15,26),p=(4,4)
95.90%
96.22%
99.13%
0.33
56.08%
66.18%
BASA
w=(10,10),p=(4,4)
95.98%
96.29%
99.17%
0.34
56.87%
67.02%
Table 4: VBench Custom evaluation on 100 extrapolated videos generated from 20 prompts with five samples per prompt. All methods extrapolate the same 81-frame videos from 832×480 to 1664×960 . SC, BC, MS, DD, AQ, and IQ denote Subject Consistency, Background Consistency, Motion Smoothness, Dynamic Degree, Aesthetic Quality, and Imaging Quality, respectively.
Method
Shift
Pool
PSNR ↑
SSIM ↑
FID ↓
LPIPS ↓
CLIP-I ↑
DINO ↑
CLIP-T ↑
IS ↑
FLUX
–
–
–
–
–
–
–
–
31.13
23.41
BASA
×
×
22.23
0.79
33.13
0.29
94.36
91.80
31.00
24.13
BASA
✓
×
23.33
0.84
24.05
0.24
96.86
94.72
31.28
23.13
BASA
✓
✓
24.09
0.86
20.97
0.20
97.60
95.95
31.28
23.36
Table 5: Ablation study for image resolution extrapolation from 1024×1024 to 2048×2048 with FLUX.1-dev. All BASA variants use w=32 and p=4 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Shift
Pool
Route
SC ↑
BC ↑
MS ↑
DD ↑
AQ ↑
IQ ↑
Wan
–
–
–
96.43%
96.84%
99.18%
0.37
57.53%
65.22%
BASA
×
✓
✓
95.06%
95.47%
99.13%
0.29
50.76%
66.06%
BASA
✓
×
✓
95.41%
96.00%
99.01%
0.32
56.20%
67.04%
BASA
✓
✓
×
95.76%
96.09%
99.17%
0.29
56.48%
65.94%
BASA
✓
✓
✓
95.98%
96.29%
99.17%
0.34
56.87%
67.02%
Appendix
Table 6: Ablation study for video resolution extrapolation with w=(10,10) and p=(4,4) . SC, BC, MS, DD, AQ, and IQ denote Subject Consistency, Background Consistency, Motion Smoothness, Dynamic Degree, Aesthetic Quality, and Imaging Quality, respectively.
Figure 5: Image extrapolation qualitative results of 1K to 2K generated by our method BASA (window size=32, pooling size=4) based on FLUX-1.dev. Left ones are 1K images generated by Original FLUX, while the right ones are extrapolation qualitative results generated by BASA
Figure 6: Video extrapolation qualitative results of 480×832 to 960×1664 generated by our method BASA (window size= 15×26 based on latents, pooling size=4) based on Wan2.1-1.3B.
Figure 7: Video extrapolation qualitative results of 480×832 to 960×1664 generated by our method BASA (window size= 10×10 based on latents, pooling size=4) based on Wan2.1-1.3B.
Figure 8: Qualitative comparison of resolution extrapolation. The left column shows the original low-resolution images generated by FLUX.1-dev and Wan-2.1 1.3B, while the right column shows the corresponding high-resolution results extrapolated by BASA with window size=32 and pooling size=4, window size=(10, 10) and pooling size=4 respectively. Orange dashed boxes highlight regions where the low-resolution inputs exhibit insufficient fine-grained details or noticeable over-smoothing. BASA produces sharper and more coherent local structures in these regions while preserving the overall image content.
Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs prohibitive computational cost due to the quadratic complexity of attention. Block sparse attention is a common approach to mitigate this by focusing computation on important regions. However, attention maps in DiTs exhibit inherently dynamic and fine-grained sparsity, which causes existing block sparse attention methods to degrade significantly in quality, especially at high sparsity ratios. In this paper, we revisit block sparse attention and derive a theoretical lower bound on attention recall to characterize the key factors governing its effectiveness. Guided by these insights, we propose DFSAttn, a training-free sparse attention framework that enables dynamic, fine-grained sparsification efficiently. DFSAttn incorporates three core designs: Hilbert curve-based token reordering to achieve fine-grained sparsity while preserving efficient GPU execution, hierarchical block scoring for accurate block importance estimation, and sparse mask caching with adaptive ratios to balance accuracy and efficiency. Experimental results demonstrate that DFSAttn consistently outperforms prior methods under high sparsity, achieving up to 2.1× end-to-end speedup while maintaining high generation quality. Our code is open-sourced and available at https://github.com/jessica-hujie/DFSAttn.
While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing sparse methods face a dilemma: dynamic pruning suffers from prohibitive runtime overhead and memory fragmentation, while static heuristics fail to capture fine-grained dependencies. In this work, we propose ScalingAttention, a training-free framework grounded in a key inductive bias: while individual activations are input-dependent, the high-mass attention regions for each head rapidly converge to a stable, prompt-agnostic Intrinsic Sparse Topology. This topology is weight-encoded, scale-invariant, and efficient to extract. ScalingAttention decouples topology discovery from sparsity control via: (1) WEST (Weight-Encoded Sparse Topology), which extracts a robust block-sparse prior mask offline to eliminate runtime search; (2) FAST (Fidelity-Aware Sensitivity Tuning), which adaptively tunes head-wise sparsity based on diffusion fidelity requirements. To ensure practical acceleration, we co-design a hardware-aligned bit-wise block-sparse kernel. Experiments on Wan2.1 show up to 1.90X end-to-end speedup with superior fidelity, establishing a new Pareto frontier over state-of-the-art baselines.
Ruiliang Zhou, Xuecheng Wu, Kang He +6
KlingAI Research · Beijing Institute of Technology · NVIDIA +1
Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57× and 1.37× end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.