Native 4K (2176×3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed T3 (Transform Trained Transformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, T3-Video introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on 4K-VBench show that T3-Video substantially outperforms existing approaches: while delivering performance improvements (+4.29↑ VQA and +0.08↑ VTC), it accelerates native 4K video generation by more than 10×. Project page at https://zhangzjn.github.io/projects/T3-Video
Figures & tables
Figure 1 : 4K (2176 × 3840) inference visualization: Wan2.1-T2V-1.3B [ 48 ] (81 f ), HunyuanVideo [ 24 ] (41 f ), UltraGen [ 20 ] (29 f ), and our T3-Video-T2V-1.3B (81 f ). Efficiency tests are performed on 81 frames with FlashAttention2 [ 8 ] on a single H20 GPU. Bar chart: blue denotes theoretical MACs while orange denotes measured latency of DiT in Wan2.1-1.3B [ 48 ] . Vertical axis is on a logarithmic scale.
Figure 2 : 720P results w/ or w/o finetuning for close/remote window-attention and T3 module by 4 × 4 blocks.
Figure 3 : Intuitive diagram for T3 strategy . Taking the typical 2D (1,1) position F1,1 of the input latent feature with a window size of 2 as an intuitive example, it uses shared attention parameters to perform information exchange across multiple scales ( Δh1/Δw1,Δh2/Δw2,⋯ ) simultaneously.
Figure 4 : T3-Video restores full-attention capability through the re-transform attention process, which is potentially applicable to efficient pre-training of new architectures.
Figure 5 : Directly scaling T3-Video-1.3B from 720P leads to performance degradation, which cannot adapt to arbitrary resolution inference. For I2V, the degradation is alleviated due to image prior.
Resolution
Encoder
Decoder
DiT
Rest
QKV proj.
O proj.
FFN
0 0 0 Attn
All
Param.
0 0 53.6M
0 0 73.3M
309.6M
212.5M
70.8M
826.1M
0 0 0 0 0
0 1419.0M
MACs
0 480 × 832 0
0 0 81.2T
0 137.0T
0 0 4.7T
0 0 7.0T
0 2.3T
0 27.1T
0 0 0 98.9T
0 0 140.0T
0 0 0 0 4.5T ×22.0↑
0 0 0 45.5T
0 720 × 1280
0 187.5T
0 316.3T
0 10.8T
0 16.1T
0 5.4T
0 62.4T
0 0 526.7T
0 0 621.4T
0 0 0 17.0T ×30.9↑
0 0 111.7T
Table 1 : Disassembly of parameters and MACs between Wan2.1-T2V-1.3B (Top) and T3-Video-T2V-1.3B (Bottom) that focuses on the last two columns (“Attn" and “ALL") of each row. Both of them contain same parameters. “Rest" includes text and time-related embeddings and cross-attention.
Resolution
T2V-1.3B
T2V-1.3B-Deployment
Decoder
DiT (50)
Latency
eVAE
DiT (8)
Latency
0 480 × 832 0
0 0 5.8
131.1
0 0 267.9
-
-
-
0 0 5.8
50.4 ×2.6↑
0 0 106.6
0.233 ×24.9↑
4.0 ×12.5↑
4.3 ×25.0↑
0 720 × 1280 0
0 13.8
572.1
0 1,157.9
-
-
-
0 13.8
123.0 ×4.7↑
0 0 259.9
0.514 ×26.9↑
9.8 ×12.5↑
10.4 ×25.1↑
1088 × 1920
0 49.8
2,653.7
0 5,357.2
-
-
-
Table 2 : Inference latency analysis of T3-Video-1.3B (denoising 50 steps with CFG) and deployment version described in Sec. 3.5 (denoising 8 steps without CFG and along with e VAE). Unit: s. Top: Official. Bottom: T3-Video. ×↑ and ×↑ denote the relative speedups in the vertical and horizontal directions, respectively.
Reso.
Train
Test
720P
29.6G
8.8G
1080P
52.8G
18.6G
4K
179.4G
59.5G
Table 3 : Memory analysis.
Figure 6 : Directly fine-tuning T3-Video-T2V-1.3B for native 4K generation: as training progresses, the model first adapts to the spatial structure and then progressively refines the fine details.
VAE
Encoder
Decoder
PSNR
SSIM
LPIPS
Params.
MACs
Params.
MAC
Latency
Speedup
Wan2.1-1.3B
0 53.60M
187.49T
0 73.30M
316.26T
13.8380
0 1.0 ×
38.07
0.9576
0.0251
e VAE-Wan2.1-1.3B-10M
0 0 1.47M
0 0 5.86T
0 0 9.84M
0 13.18T
0 0.5145
26.9 ×
36.29
0.9422
0.04
Wan2.2-5B
149.64M
130.82T
555.05M
688.58T
10.5796
0 1.0 ×
38.30
0.9567
0.0324
e VAE-Wan2.2-5B-35M
149.64M
130.82T
0 34.97M
0 43.34T
0 1.3040
0 8.1 ×
37.14
0.9484
0.052
Table 4 : Efficiency and performance of efficient e VAE over official Wan2.1-1.3-VAE and Wan2.2-5B-VAE for faster inference. Defalut 720 × 1280 resolution on one H20 GPU.
Figure 7 : Formidable direct LoRA may fail.
Model
VQA
VTC
DoG
BM
RA
TDS
TEP
Wan2.1-T2V-1.3B (Official)
30.01
0.37
0.22
0.968
0.76
0.83
0.33
HunyuanVideo (Official)
61.92
0.68
0.26
0.961
0.72
0.91
0.34
UltraGen [ 20 ]
67.43
0.75
0.45
0.980
0.72
0.84
0.52
T3-Video-T2V-1.3B (Ours)
71.72
0.83
0.54
0.988
0.83
0.91
0.70
T3-Video-T2V-1.3B-LoRA (Ours)
70.78
0.79
0.50
0.990
0.79
0.91
0.61
Table 5 : Comparison with SoTAs on native 4K Video generation with pretrained models from UltraGen [ 20 ] .
Model
VQA
VTC
DoG
BM
RA
TDS
TEP
I2V
Wan2.1-I2V-1.3B (Official)
59.75
0.79
0.27
0.974
0.65
0.82
0.32
T3-Video-I2V-1.3B (Ours)
63.6
0.82
0.32
0.986
0.83
0.92
0.41
T2V
Wan2.2-T2V-5B (Official)
47.23
0.52
0.16
0.984
0.51
0.85
0.25
T3-Video-T2V-5B (Ours)
67.40
0.91
0.35
0.995
0.82
0.88
0.34
I2V
Wan2.2-I2V-5B (Official)
65.13
0.89
0.28
0.989
0.67
0.86
0.38
T3-Video-I2V-5B (Ours)
68.84
0.90
0.34
0.996
0.84
0.90
0.44
Table 6 : Multi-baseline generalization of T3-Video (4K).
Figure 8 : T3-Video series (T2V and I2V) based on Wan2.1-1.3B and Wan2.2-5B achieve satisfactory results in both full-/LoRA-tuning.
Model
VQA
VTC
DoG
BM
RA
TDS
TEP
(a) Return Official
Wan2.1-T2V-1.3B (Official)
70.56
0.91
0.42
0.938
0.71
0.87
0.51
T3-Video-T2V-1.3B (Ours)
69.37
0.90
0.40
0.948
0.72
0.89
0.49
Wan2.1-T2V-1.3B (Return Ours)
69.51
0.90
0.42
0.925
0.71
0.88
0.49
(b) Batch Size
8
66.65
0.81
0.30
0.928
0.67
0.87
0.47
16
67.17
0.83
0.39
0.924
0.66
0.88
0.42
32
68.05
0.84
0.36
0.920
0.71
0.88
0.49
Table 7 : Empirical observations on basic factors (720P).
Models
Subject Consistency
Background Consistency
Temporal Flickering
Motion Smoothness
Dynamic Degree
Aesthetic Quality
Imaging Quality
Object Class
UltraWAN-4K (LoRA, 29f)
96.05%
98.02%
98.88%
98.47% ∗
66.66% ∗
56.81%
71.61%
50.00%
T3-Video-4K (FT, 81f)
97.17%
98.10%
98.52%
98.85%*
66.66%*
59.83%
72.18%
66.66%
Models
Multiple Objects
Human Action
Color
Spatial Relationship
Scene
Appearance Style
Temporal Style
Overall Consistency
UltraWAN-4K (LoRA, 29f)
42.75%
66.66%
100.0%
100.0%
00.00%
19.46%
19.31%
22.88%
T3-Video-4K (FT, 81f)
45.62%
66.66%
100.0%
100.0%
16.66%
19.28%
21.15%
23.61%
Table 8 : VBench evaluation results per dimension. ∗ : Videos are downsampled to 1K to avoid OOM.
Method
➀
➁
➂
➃
UltraGen [ 20 ]
28.75%
40.25%
43.08%
36.58%
Ours
71.25%
59.75%
56.92%
63.42%
Table 9 : Human study with UltraGen [ 20 ] . ➀ Video Quality, ➁ Text Consistency, ➂ Temporal Consistency, and ➃ Detail Richness.
Recent progress in transformer-based architectures has demonstrated remarkable success in video generation tasks. However, the quadratic complexity of full attention mechanisms remains a critical bottleneck, particularly for high-resolution and long-duration video sequences. In this paper, we propose NABLA, a novel Neighborhood Adaptive Block-Level Attention mechanism that dynamically adapts to sparsity patterns in video diffusion transformers (DiTs). By leveraging block-wise attention with adaptive sparsity-driven threshold, NABLA reduces computational overhead while preserving generative quality. Our method does not require custom low-level operator design and can be seamlessly integrated with PyTorch's Flex Attention operator. Experiments demonstrate that NABLA achieves up to 2.7x faster training and inference compared to baseline almost without compromising quantitative metrics (CLIP score, VBench score, human evaluation score) and visual quality drop. The code and model weights are available here: https://github.com/gen-ai-team/Wan2.1-NABLA
Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva +6
Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs prohibitive computational cost due to the quadratic complexity of attention. Block sparse attention is a common approach to mitigate this by focusing computation on important regions. However, attention maps in DiTs exhibit inherently dynamic and fine-grained sparsity, which causes existing block sparse attention methods to degrade significantly in quality, especially at high sparsity ratios. In this paper, we revisit block sparse attention and derive a theoretical lower bound on attention recall to characterize the key factors governing its effectiveness. Guided by these insights, we propose DFSAttn, a training-free sparse attention framework that enables dynamic, fine-grained sparsification efficiently. DFSAttn incorporates three core designs: Hilbert curve-based token reordering to achieve fine-grained sparsity while preserving efficient GPU execution, hierarchical block scoring for accurate block importance estimation, and sparse mask caching with adaptive ratios to balance accuracy and efficiency. Experimental results demonstrate that DFSAttn consistently outperforms prior methods under high sparsity, achieving up to 2.1× end-to-end speedup while maintaining high generation quality. Our code is open-sourced and available at https://github.com/jessica-hujie/DFSAttn.
Diffusion Transformers achieve strong video generation quality, but the quadratic cost of full attention limits efficiency. We introduce OSP-Next, an efficient text-to-video generation model that integrates sparse attention, parallelism, quantization, and reinforcement learning. OSP-Next uses a hybrid full-sparse attention architecture, where the sparse component is implemented with Skiparse-2D Attention. This fixed-pattern mechanism applies token-wise and group-wise sparse attention along spatial dimensions, leveraging locality while maintaining native compatibility with FlashAttention kernels. Based on the local equivalence of rearrangement in Skiparse-2D Attention, we further propose Sparse Sequence Parallelism (SSP), which partitions subsequences across ranks and switches sparse patterns through a single All-to-All communication. Compared with Ulysses Sequence Parallelism (SP), SSP provides a native parallel strategy for sparse attention and reduces communication volume by 75%. OSP-Next also incorporates HiF8 quantization to enable stable joint training with 8-bit quantization and sparse fine-tuning, and applies Mix-GRPO post-training to improve the performance of the sparse model. Experiments show that OSP-Next achieves a VBench total score of 83.73%, surpassing the Wan2.1 baseline. Under the 5-second 720P and 5-second 768P settings, OSP-Next achieves up to 1.64× single-GPU speedup and over 1.52× eight-GPU speedup on NVIDIA H200 GPUs. In addition, with only a 0.4% drop in VBench total score, OSP-Next-HiF8 achieves 1.69× and 2.27× speedups under the two settings on a single Ascend 950PR, demonstrating the efficiency and performance of OSP-Next across hardware platforms.
Yunyang Ge, Xianyi He, Zezhong Zhang +4
1Peking University · 2Nanyang Technological University, Singapore · 3Rabbitpre AI