Native 4K (2176×3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed T3 (Transform Trained Transformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, T3-Video introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on 4K-VBench show that T3-Video substantially outperforms existing approaches: while delivering performance improvements (+4.29↑ VQA and +0.08↑ VTC), it accelerates native 4K video generation by more than 10×. Project page at https://zhangzjn.github.io/projects/T3-Video
Figures & tables
Figure 1 : 4K (2176 × 3840) inference visualization: Wan2.1-T2V-1.3B [ 48 ] (81 f ), HunyuanVideo [ 24 ] (41 f ), UltraGen [ 20 ] (29 f ), and our T3-Video-T2V-1.3B (81 f ). Efficiency tests are performed on 81 frames with FlashAttention2 [ 8 ] on a single H20 GPU. Bar chart: blue denotes theoretical MACs while orange denotes measured latency of DiT in Wan2.1-1.3B [ 48 ] . Vertical axis is on a logarithmic scale.
Figure 2 : 720P results w/ or w/o finetuning for close/remote window-attention and T3 module by 4 × 4 blocks.
Figure 3 : Intuitive diagram for T3 strategy . Taking the typical 2D (1,1) position F1,1 of the input latent feature with a window size of 2 as an intuitive example, it uses shared attention parameters to perform information exchange across multiple scales ( Δh1/Δw1,Δh2/Δw2,⋯ ) simultaneously.
Figure 4 : T3-Video restores full-attention capability through the re-transform attention process, which is potentially applicable to efficient pre-training of new architectures.
Figure 5 : Directly scaling T3-Video-1.3B from 720P leads to performance degradation, which cannot adapt to arbitrary resolution inference. For I2V, the degradation is alleviated due to image prior.
Resolution
Encoder
Decoder
DiT
Rest
QKV proj.
O proj.
FFN
0 0 0 Attn
All
Param.
0 0 53.6M
0 0 73.3M
309.6M
212.5M
70.8M
826.1M
0 0 0 0 0
0 1419.0M
MACs
0 480 × 832 0
0 0 81.2T
0 137.0T
0 0 4.7T
0 0 7.0T
0 2.3T
0 27.1T
0 0 0 98.9T
0 0 140.0T
0 0 0 0 4.5T ×22.0↑
0 0 0 45.5T
0 720 × 1280
0 187.5T
0 316.3T
0 10.8T
0 16.1T
0 5.4T
0 62.4T
0 0 526.7T
0 0 621.4T
0 0 0 17.0T ×30.9↑
0 0 111.7T
Table 1 : Disassembly of parameters and MACs between Wan2.1-T2V-1.3B (Top) and T3-Video-T2V-1.3B (Bottom) that focuses on the last two columns (“Attn" and “ALL") of each row. Both of them contain same parameters. “Rest" includes text and time-related embeddings and cross-attention.
Resolution
T2V-1.3B
T2V-1.3B-Deployment
Decoder
DiT (50)
Latency
eVAE
DiT (8)
Latency
0 480 × 832 0
0 0 5.8
131.1
0 0 267.9
-
-
-
0 0 5.8
50.4 ×2.6↑
0 0 106.6
0.233 ×24.9↑
4.0 ×12.5↑
4.3 ×25.0↑
0 720 × 1280 0
0 13.8
572.1
0 1,157.9
-
-
-
0 13.8
123.0 ×4.7↑
0 0 259.9
0.514 ×26.9↑
9.8 ×12.5↑
10.4 ×25.1↑
1088 × 1920
0 49.8
2,653.7
0 5,357.2
-
-
-
Table 2 : Inference latency analysis of T3-Video-1.3B (denoising 50 steps with CFG) and deployment version described in Sec. 3.5 (denoising 8 steps without CFG and along with e VAE). Unit: s. Top: Official. Bottom: T3-Video. ×↑ and ×↑ denote the relative speedups in the vertical and horizontal directions, respectively.
Reso.
Train
Test
720P
29.6G
8.8G
1080P
52.8G
18.6G
4K
179.4G
59.5G
Table 3 : Memory analysis.
Figure 6 : Directly fine-tuning T3-Video-T2V-1.3B for native 4K generation: as training progresses, the model first adapts to the spatial structure and then progressively refines the fine details.
VAE
Encoder
Decoder
PSNR
SSIM
LPIPS
Params.
MACs
Params.
MAC
Latency
Speedup
Wan2.1-1.3B
0 53.60M
187.49T
0 73.30M
316.26T
13.8380
0 1.0 ×
38.07
0.9576
0.0251
e VAE-Wan2.1-1.3B-10M
0 0 1.47M
0 0 5.86T
0 0 9.84M
0 13.18T
0 0.5145
26.9 ×
36.29
0.9422
0.04
Wan2.2-5B
149.64M
130.82T
555.05M
688.58T
10.5796
0 1.0 ×
38.30
0.9567
0.0324
e VAE-Wan2.2-5B-35M
149.64M
130.82T
0 34.97M
0 43.34T
0 1.3040
0 8.1 ×
37.14
0.9484
0.052
Table 4 : Efficiency and performance of efficient e VAE over official Wan2.1-1.3-VAE and Wan2.2-5B-VAE for faster inference. Defalut 720 × 1280 resolution on one H20 GPU.
Figure 7 : Formidable direct LoRA may fail.
Model
VQA
VTC
DoG
BM
RA
TDS
TEP
Wan2.1-T2V-1.3B (Official)
30.01
0.37
0.22
0.968
0.76
0.83
0.33
HunyuanVideo (Official)
61.92
0.68
0.26
0.961
0.72
0.91
0.34
UltraGen [ 20 ]
67.43
0.75
0.45
0.980
0.72
0.84
0.52
T3-Video-T2V-1.3B (Ours)
71.72
0.83
0.54
0.988
0.83
0.91
0.70
T3-Video-T2V-1.3B-LoRA (Ours)
70.78
0.79
0.50
0.990
0.79
0.91
0.61
Table 5 : Comparison with SoTAs on native 4K Video generation with pretrained models from UltraGen [ 20 ] .
Model
VQA
VTC
DoG
BM
RA
TDS
TEP
I2V
Wan2.1-I2V-1.3B (Official)
59.75
0.79
0.27
0.974
0.65
0.82
0.32
T3-Video-I2V-1.3B (Ours)
63.6
0.82
0.32
0.986
0.83
0.92
0.41
T2V
Wan2.2-T2V-5B (Official)
47.23
0.52
0.16
0.984
0.51
0.85
0.25
T3-Video-T2V-5B (Ours)
67.40
0.91
0.35
0.995
0.82
0.88
0.34
I2V
Wan2.2-I2V-5B (Official)
65.13
0.89
0.28
0.989
0.67
0.86
0.38
T3-Video-I2V-5B (Ours)
68.84
0.90
0.34
0.996
0.84
0.90
0.44
Table 6 : Multi-baseline generalization of T3-Video (4K).
Figure 8 : T3-Video series (T2V and I2V) based on Wan2.1-1.3B and Wan2.2-5B achieve satisfactory results in both full-/LoRA-tuning.
Model
VQA
VTC
DoG
BM
RA
TDS
TEP
(a) Return Official
Wan2.1-T2V-1.3B (Official)
70.56
0.91
0.42
0.938
0.71
0.87
0.51
T3-Video-T2V-1.3B (Ours)
69.37
0.90
0.40
0.948
0.72
0.89
0.49
Wan2.1-T2V-1.3B (Return Ours)
69.51
0.90
0.42
0.925
0.71
0.88
0.49
(b) Batch Size
8
66.65
0.81
0.30
0.928
0.67
0.87
0.47
16
67.17
0.83
0.39
0.924
0.66
0.88
0.42
32
68.05
0.84
0.36
0.920
0.71
0.88
0.49
Table 7 : Empirical observations on basic factors (720P).
Models
Subject Consistency
Background Consistency
Temporal Flickering
Motion Smoothness
Dynamic Degree
Aesthetic Quality
Imaging Quality
Object Class
UltraWAN-4K (LoRA, 29f)
96.05%
98.02%
98.88%
98.47% ∗
66.66% ∗
56.81%
71.61%
50.00%
T3-Video-4K (FT, 81f)
97.17%
98.10%
98.52%
98.85%*
66.66%*
59.83%
72.18%
66.66%
Models
Multiple Objects
Human Action
Color
Spatial Relationship
Scene
Appearance Style
Temporal Style
Overall Consistency
UltraWAN-4K (LoRA, 29f)
42.75%
66.66%
100.0%
100.0%
00.00%
19.46%
19.31%
22.88%
T3-Video-4K (FT, 81f)
45.62%
66.66%
100.0%
100.0%
16.66%
19.28%
21.15%
23.61%
Table 8 : VBench evaluation results per dimension. ∗ : Videos are downsampled to 1K to avoid OOM.
Method
➀
➁
➂
➃
UltraGen [ 20 ]
28.75%
40.25%
43.08%
36.58%
Ours
71.25%
59.75%
56.92%
63.42%
Table 9 : Human study with UltraGen [ 20 ] . ➀ Video Quality, ➁ Text Consistency, ➂ Temporal Consistency, and ➃ Detail Richness.