cs.CVSep 28, 2026

TSGate: Timestep-Aware Gated Attention for Diffusion Transformers

Authors: Boyu Zhang, Yifan Liu, Shuxia Lin, Qingjian Ni, Yinfei Xu, Xu Yang

Organizations: Alibaba Token Hub, Alibaba Group · Southeast University

Abstract

Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-of-domain (OOD) prompts, including the free-form descriptions supplied by users at inference time. Although LLM-based rewriting can convert these prompts into structured formats, it does not guarantee that the rewritten prompts align with the training distribution. Our analysis links this degradation to attention sinks and reduced early-step image-to-text attention and shows that sink suppression alone is insufficient to restore generation quality. Despite effective sink suppression, models trained with standard gated attention exhibit reduced early-step image-to-text attention and suboptimal generation quality. Based on these insights, we propose Timestep-Aware Gated Attention (TSGate), which injects a timestep-conditioned bias into the gate signal so that gating behavior adapts across denoising steps. Extensive experiments show that TSGate consistently outperforms both the baseline and standard gated attention across multiple benchmarks, improving the raw-prompt DPG score by 9.5% over the baseline.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers

    May 21, 2026Javad Rajabi, Kimia Shaban, Koorosh Roohi +2Diffusion TransformersRotary Position Embedding

  2. AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation

    Mar 13, 2026Xuanhua Yin, Chuanzhi Xu, Haoxian Zhou +2Diffusion TransformersImage Generation

  3. Denoising Diffusion Generative Models Secretly Calculate Attentions

    Sep 1, 2026Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3Denoising Diffusion Probabilistic ModelImage Generation