Organizations: School of Artificial Intelligence, Shanghai Jiao Tong University · School of Data Science, Fudan University · School of Computing and Data Science, The University of Hong Kong
Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity. However, isolated windows block cross-window interaction, often introducing visible grid-like artifacts in the generated results. Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality. However, its irregular computation patterns create a substantial gap between theoretical and practical speedups and require specialized kernels tailored to each hardware backend. To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration. Specifically, BASA replaces visual self-attention with shifted local-window attention. By introducing a structured window-shifting scheme across DiT blocks, we allow tokens divided by window boundaries in one layer to communicate in the following layers, thereby achieving global information exchange and eliminating window-induced visual artifacts. Notably, our design introduces no additional irregular operators or customized kernels, making it readily deployable on existing attention backends and closing the gap between theoretical sparsity and practical acceleration. Experiments demonstrate that BASA achieves measured speedups exceeding 90% of the theoretical estimates on FLUX and delivers a 4.52× attention speedup on Wan while maintaining competitive generation quality. Codes are publicly available at: https://github.com/lama0110/BASA.
Figures & tables
Figure 1: Ultra-resolution results generated by FLUX.1-dev and Wan2.1-1.3B with our approach. Resolution is marked on the bottom-right corner of each result in the format of width×height. Corresponding prompts can be found in the appendix.
Figure 2: Motivation for BASA. Swin-style window attention ( Liu et al., 2021 ) provides fast inference but introduces severe window boundary artifacts. CLEAR ( Liu et al., 2025 ) preserves high visual fidelity, while its practical speed gain is relatively limited. BASA preserves faithful reconstruction while achieving efficient inference, providing a favorable quality–efficiency trade-off.
Figure 3: Overview of BASA. (a) The proposed sparse attention framework integrates multiple components to recover both local and global visual interactions under reduced attention cost. (b) BASA is applied as a drop-in replacement for the visual self-attention module in DiT blocks, while the remaining generation pipeline remains unchanged.
Figure 4: Qualitative comparison of image resolution extrapolation from 1024×1024 to 2048×2048 on FLUX.1-dev. Competing sparse-attention methods may exhibit window-boundary artifacts or duplicated structures in some cases, while BASA maintains or adds more consistent spatial details.
Method
Config
PSNR ↑
SSIM ↑
FID ↓
LPIPS ↓
CLIP-I ↑
DINO ↑
CLIP-T ↑
IS ↑
FLUX
–
–
–
–
–
–
–
31.13
23.41
CLEAR
r=16 , distilled
23.50
0.83
23.28
0.23
97.09
94.94
31.16
23.44
Swin
w=32 , distilled
23.33
0.83
34.10
0.28
93.62
91.62
31.43
23.67
STA
w=32 , distilled
23.51
0.83
24.02
0.24
96.45
94.57
31.29
23.68
BASA
w=8,p=4 , distilled
21.32
0.78
38.65
0.31
92.61
87.59
31.48
22.50
BASA
w=16,p=4 , distilled
22.57
0.82
26.67
0.25
96.45
93.53
31.34
24.10
Table 1: Quantitative comparison for image resolution extrapolation from 1024×1024 to 2048×2048 with FLUX.1-dev. PSNR, SSIM, FID, LPIPS, CLIP-I, and DINO are computed against the corresponding dense-FLUX output.
Method
Config.
Implementation
Latency (ms) ↓
Act. ↑
Theo. ↑
Real. ↑
FLUX
Dense
SDPA
21.666
1.000
1.000
1.000
Swin
w=32
SDPA
8.819
2.455
3.187
0.770
CLEAR
r=16
FlexAttn
13.106
1.629
3.229
0.504
STA
w=32
Triton
12.955
1.648
3.092
0.533
BASA
w=8,p=4
SDPA
8.079
2.680
3.064
0.874
BASA
w=16,p=4
SDPA
8.131
2.665
2.987
0.892
Table 2: Theoretical and measured weighted-MSA efficiency for 1024×1024 to 2048×2048 resolution extrapolation. Weighted MSA denotes the average full multi-head self-attention latency per layer, weighted over all double- and single-stream attention layers of FLUX and the corresponding window schedules. Backend indicates the attention execution backend used in our implementation. Act., Theo., and Real. denote measured speedup, theoretical speedup, and their ratio, respectively.
Method
Config (latent)
Impl.
Sparsity
TFLOPs (Attn.)
Latency (ms)
MFU
Kernel Eff.
Speedup
Wan
–
SDPA
0.00%
106.12
509.66
66.74%
100.00%
1.00×
CLEAR
w=16
Flex
89.46%
11.18
485.79
7.38%
11.05%
1.05×
Swin
w=16
Flex
96.25%
3.84
396.60
3.10%
4.65%
1.29×
STA
w=(21,30,39),t=(3,10,13)
Triton
81.25%
19.90
175.09
36.42%
53.48%
2.85×
BASA
w=(10,10),p=(2,2)
SDPA
70.17%
31.66
255.87
39.66%
59.42%
1.99×
BASA
w=(15,26),p=(4,4)
SDPA
81.25%
19.90
126.07
50.59%
75.80%
4.04×
Table 3: Attention efficiency for video resolution extrapolation from 832×480 to 1664×960 with Wan2.1-T2V-1.3B. Configurations are specified on the latent grid. Latency measures attention-only forward time after warm-up, and speedup is reported relative to dense Wan.
Method
Config
SC ↑
BC ↑
MS ↑
DD ↑
AQ ↑
IQ ↑
Wan
–
96.43%
96.84%
99.18%
0.37
57.53%
65.22%
CLEAR
w=(8,8),d=4
96.27%
96.54%
99.19%
0.35
57.50%
67.11%
Swin
w=(8,8)
93.62%
94.59%
98.15%
0.04
41.36%
60.90%
STA
w=(18,24,24),t=(6,8,8)
95.89%
96.05%
99.16%
0.37
56.26%
67.56%
BASA
w=(15,26),p=(4,4)
95.90%
96.22%
99.13%
0.33
56.08%
66.18%
BASA
w=(10,10),p=(4,4)
95.98%
96.29%
99.17%
0.34
56.87%
67.02%
Table 4: VBench Custom evaluation on 100 extrapolated videos generated from 20 prompts with five samples per prompt. All methods extrapolate the same 81-frame videos from 832×480 to 1664×960 . SC, BC, MS, DD, AQ, and IQ denote Subject Consistency, Background Consistency, Motion Smoothness, Dynamic Degree, Aesthetic Quality, and Imaging Quality, respectively.
Method
Shift
Pool
PSNR ↑
SSIM ↑
FID ↓
LPIPS ↓
CLIP-I ↑
DINO ↑
CLIP-T ↑
IS ↑
FLUX
–
–
–
–
–
–
–
–
31.13
23.41
BASA
×
×
22.23
0.79
33.13
0.29
94.36
91.80
31.00
24.13
BASA
✓
×
23.33
0.84
24.05
0.24
96.86
94.72
31.28
23.13
BASA
✓
✓
24.09
0.86
20.97
0.20
97.60
95.95
31.28
23.36
Table 5: Ablation study for image resolution extrapolation from 1024×1024 to 2048×2048 with FLUX.1-dev. All BASA variants use w=32 and p=4 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Shift
Pool
Route
SC ↑
BC ↑
MS ↑
DD ↑
AQ ↑
IQ ↑
Wan
–
–
–
96.43%
96.84%
99.18%
0.37
57.53%
65.22%
BASA
×
✓
✓
95.06%
95.47%
99.13%
0.29
50.76%
66.06%
BASA
✓
×
✓
95.41%
96.00%
99.01%
0.32
56.20%
67.04%
BASA
✓
✓
×
95.76%
96.09%
99.17%
0.29
56.48%
65.94%
BASA
✓
✓
✓
95.98%
96.29%
99.17%
0.34
56.87%
67.02%
Appendix
Table 6: Ablation study for video resolution extrapolation with w=(10,10) and p=(4,4) . SC, BC, MS, DD, AQ, and IQ denote Subject Consistency, Background Consistency, Motion Smoothness, Dynamic Degree, Aesthetic Quality, and Imaging Quality, respectively.
Figure 5: Image extrapolation qualitative results of 1K to 2K generated by our method BASA (window size=32, pooling size=4) based on FLUX-1.dev. Left ones are 1K images generated by Original FLUX, while the right ones are extrapolation qualitative results generated by BASA
Figure 6: Video extrapolation qualitative results of 480×832 to 960×1664 generated by our method BASA (window size= 15×26 based on latents, pooling size=4) based on Wan2.1-1.3B.
Figure 7: Video extrapolation qualitative results of 480×832 to 960×1664 generated by our method BASA (window size= 10×10 based on latents, pooling size=4) based on Wan2.1-1.3B.
Figure 8: Qualitative comparison of resolution extrapolation. The left column shows the original low-resolution images generated by FLUX.1-dev and Wan-2.1 1.3B, while the right column shows the corresponding high-resolution results extrapolated by BASA with window size=32 and pooling size=4, window size=(10, 10) and pooling size=4 respectively. Orange dashed boxes highlight regions where the low-resolution inputs exhibit insufficient fine-grained details or noticeable over-smoothing. BASA produces sharper and more coherent local structures in these regions while preserving the overall image content.