cs.CVDec 15, 2025

Transform Trained Transformer for Accelerating Native 4K Video Generation

Authors: Jiangning Zhang, Junwei Zhu, Teng Hu, Yabiao Wang, Donghao Luo, Weijian Cao, Zhenye Gan, Xiaobin Hu, +4 more

Organizations: Youtu Lab, Tencent · Zhejiang University

Abstract

Native 4K (2176×\times3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed T3 (T\textbf{T}ransform T\textbf{T}rained T\textbf{T}ransformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, T3-Video\textbf{T3-Video} introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on 4K-VBench\textbf{4K-VBench} show that T3-Video\textbf{T3-Video} substantially outperforms existing approaches: while delivering performance improvements (+4.29↑\uparrow VQA and +0.08↑\uparrow VTC), it accelerates native 4K video generation by more than 10×\times. Project page at https://zhangzjn.github.io/projects/T3-Video

Figures & tables

Explore similar work

CardsList
  1. NABLA: Neighborhood Adaptive Block-Level Attention

    Jul 17, 2025Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva +6Generative QualitySparsity

  2. OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning

    May 27, 2026Yunyang Ge, Xianyi He, Zezhong Zhang +4Text-To-Video Generation ModelNext