Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-of-domain (OOD) prompts, including the free-form descriptions supplied by users at inference time. Although LLM-based rewriting can convert these prompts into structured formats, it does not guarantee that the rewritten prompts align with the training distribution. Our analysis links this degradation to attention sinks and reduced early-step image-to-text attention and shows that sink suppression alone is insufficient to restore generation quality. Despite effective sink suppression, models trained with standard gated attention exhibit reduced early-step image-to-text attention and suboptimal generation quality. Based on these insights, we propose Timestep-Aware Gated Attention (TSGate), which injects a timestep-conditioned bias into the gate signal so that gating behavior adapts across denoising steps. Extensive experiments show that TSGate consistently outperforms both the baseline and standard gated attention across multiple benchmarks, improving the raw-prompt DPG score by 9.5% over the baseline.
Figures & tables
Figure 1: Sink suppression is not enough; early text interaction matters. For two out-of-domain prompts, TSGate renders the requested snow-capped peak and parrot’s head more clearly than the baseline and standard gated attention. The schematic curves and bars illustrate the motivation for timestep-aware gating: reducing attention sinks must be accompanied by stronger text guidance during early denoising.
Figure 2: Overview of TSGate. At each denoising step t , a channel-wise time bias Bts(t) is shared across tokens and added to the content logits C ; a sigmoid produces the gate G . The gate scales the attention features before the output projection.
Figure 3: Paired generations from in-domain structured prompts and OOD prompts. Across three representative cases, departing from the training-time template substantially degrades object fidelity, compositionality, and scene coherence.
Figure 4: Attention sinks persist across depth and denoising time. Image-to-text attention maps for an in-domain example at two layers and two denoising steps. Bright vertical bands show attention concentrated on the same text tokens across image-token positions.
Figure 5: Image-to-text attention across denoising for in-domain and OOD prompts. Across five layers and three paired cases, OOD prompts (dashed curves) exhibit an early-step attention deficit relative to in-domain prompts (solid curves).
Model
DPG raw ↑
DPG full recaption ↑
GenEval ↑
CompBench ↑
Baseline
55.483
78.637
0.7861
0.5224
Gated attention
58.976
78.327
0.7862
0.5155
TSGate
60.749
79.355
0.8064
0.5236
Table 1: TSGate improves generation quality across prompt formats and benchmarks. We compare Baseline, Gated attention, and TSGate on DPG-Bench with raw and full-recaption prompts, and on GenEval and CompBench. Higher is better; DPG scores use a 0–100 scale, while the others use 0–1. Bold marks the best result among the three models shown; complete results and confidence intervals for all five variants appear in Appendix C .
Figure 6: TSGate exhibits stronger early text attention and lower late text attention than Baseline. Panels show Layer 5 image-to-text attention for all prompts, short prompts, and long prompts, from left to right. Shaded bands show 95% confidence intervals. Red annotations report the early TSGate–Gated-attention gap in percentage points (pp), with the relative increase in parentheses.
Figure 7: TSGate is more robust as prompt structure is removed. Left: DPG-Bench scores and the TSGate–Baseline gap across six prompt conditions, from the structured training format to the raw OOD prompt. Right: examples under the same conditions, with TSGate above Baseline.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Baseline
Gated attention
TSGate
TSGate- L0
TSGate- head
Transformer layers
12
12
12
12
12
Hidden width D
1,536
1,536
1,536
1,536
1,536
Attention heads
12
12
12
12
12
Head dimension
128
128
128
128
128
Text input width
3,584
3,584
3,584
3,584
3,584
Routed experts E
8
8
8
8
8
Appendix
Table 2: Architecture and gate settings of the five MoE configurations. Gated attention ( Qiu et al., 2025 ) uses content-only gating; TSGate-L0 adds the time branch only at Layer 0. EW and HW denote element-wise and head-wise gating. Parameter counts are in millions (M) and include all saved streams; totals are rounded to 0.001M.
Configuration
Baseline
TSGate
Transformer layers
24
24
Hidden width D
1,536
1,536
Attention heads
12
12
Head dimension
128
128
Text input width
4,096
4,096
Image FFN width
3,072
3,072
Appendix
Table 3: Dense backbone configurations for the ungated baseline and TSGate. All 24 image FFNs are dense, so routed/shared-expert settings are inactive. Audio FFN width is a configured value, although generation uses image and text streams. Parameter counts are in millions (M) and include all saved streams; totals are rounded to 0.001M.
Raw
Full recaption
Model
DPG score
DPG score
Baseline
55.483
78.637
Gated attention
58.976
78.327
TSGate
60.749
79.355
TSGate-L0
60.424
79.264
TSGate-head
55.069
78.858
Appendix
Table 4: MoE DPG-Bench scores and paired differences. Raw and full-recaption prompts use the same 1,065 sources, with scores averaged over four seeds per source. Scores use a 0–100 scale. Bold and underlining mark the highest and second-highest means in each prompt condition; intervals are pointwise 95% source-bootstrap confidence intervals (CIs).
Input
N
Baseline
TSGate
Δ
95% CI
Raw
1,064
41.619
64.159
+22.540
[21.175,23.884]
Full recaption
1,062
81.809
82.297
+0.489
[−0.084,1.068]
Appendix
Table 5: Dense DPG-Bench scores on common sources. Scores use a 0–100 scale. N counts paired sources; Δ is TSGate minus the baseline, computed before rounding. Intervals are pointwise 95% source-bootstrap CIs for the paired differences.
Raw
Full recaption
Category
Baseline
TSGate
Baseline
TSGate
L1 categories
Global
76.287
75.758
89.914
85.226
Entity
68.612
77.344
89.647
90.350
Attribute
65.061
78.597
89.015
89.309
Relation
75.489
85.862
89.395
91.396
Appendix
Table 6: Dense DPG-Bench category scores. Raw and full-recaption scores (0–100) retain the evaluator’s original aggregation and source coverage for first-level (L1) and second-level (L2) categories.
GenEval
T2I-CompBench
Model
Score
95% CI
Score
95% CI
Baseline
0.7861
[0.7583,0.8124]
0.5224
[0.5057,0.5394]
Gated attention
0.7862
[0.7598,0.8119]
0.5155
[0.4984,0.5326]
TSGate
0.8064
[0.7796,0.8318]
0.5236
[0.5070,0.5404]
TSGate-L0
0.7887
[0.7620,0.8147]
0.5231
[0.5066,0.5398]
TSGate-head
0.7972
[0.7703,0.8231]
0.5264
[0.5097,0.5432]
Appendix
Table 7: GenEval and T2I-CompBench macro scores and paired differences. Scores use a 0–1 scale. The six categories are equally weighted after averaging scores over four seeds within each source; pointwise 95% CIs use 10,000 within-category bootstrap resamples. All models use full recaption. Bold and underlining mark the highest and second-highest means for each benchmark.
Category
N
Baseline
Gated attention
TSGate
TSGate- L0
TSGate- head
Δ
Single object
80
0.9500
0.9563
0.9625
0.9563
0.9437
+0.0062
Two objects
99
0.8283
0.8308
0.8359
0.8485
0.8409
+0.0051
Counting
80
0.6281
0.6375
0.6594
0.6250
0.6531
+0.0219
Colors
94
0.8404
0.8404
0.8457
0.8351
0.8431
+0.0053
Position
100
0.8175
0.7800
0.8250
0.7850
0.8325
+0.0450
Color attribute
100
0.6525
0.6725
0.7100
0.6825
0.6700
+0.0375
Appendix
Table 8: GenEval category means. Scores use a 0–1 scale. N is the number of source prompts and Δ denotes TSGate − Gated attention, computed before rounding. Bold marks the highest mean in each row.
Category
N
Baseline
Gated attention
TSGate
TSGate- L0
TSGate- head
Δ
Color
150
0.8145
0.8092
0.8264
0.8217
0.8200
+0.0172
Shape
150
0.5912
0.5876
0.5903
0.6006
0.6027
+0.0027
Texture
150
0.6954
0.6802
0.6852
0.6998
0.6913
+0.0050
Spatial
150
0.3409
0.3193
0.3444
0.3217
0.3480
+0.0251
Non-spatial
150
0.3043
0.3033
0.3044
0.3038
0.3053
+0.0011
Complex
150
0.3879
0.3932
0.3907
0.3913
0.3910
−0.0026
Appendix
Table 9: T2I-CompBench category means. Scores use a 0–1 scale. N is the number of source prompts and Δ denotes TSGate − Gated attention, computed before rounding. Bold marks the highest mean in each row.
Figure 16: Attention sinks concentrate at Layer 0 across all three variants. Heatmaps show prompt-averaged Sink Scores across layers and denoising steps. The Sink Score is the maximum normalized column-sum of the joint-attention map; darker cells indicate stronger sinks. Attention weights are from the conditional branch and measured before the output gate.
Figure 17: Effects of time-bias interventions with model weights fixed. On 200 raw D-mech sources, points show mean DPG changes relative to native inference; whiskers show paired source-bootstrap 95% CIs. The x-axis omits (−38,−5) with equal unit spacing on the two segments; the dashed line marks zero.
Input / contrast
Intervention
DPG Δ
95% CI
Raw
Reverse
−2.614
[−4.281,−0.994]
Raw
Matched-mean
−0.929
[−2.455,0.560]
Raw
Zero
−42.831
[−46.078,−39.485]
Structured
Reverse
−0.919
[−2.854,0.948]
Structured
Matched-mean
−0.027
[−1.625,1.619]
Structured
Zero
−39.987
[−43.457,−36.513]
Appendix
Table 10: DPG changes under time-bias interventions. Δ is the score change relative to native inference. The last three rows subtract the structured-prompt change from the raw-prompt change. All rows use 200 paired sources; intervals are pointwise 95% source-bootstrap CIs.
Figure 18: Generation quality across six prompt representations. Both panels use the same 300 sources. (a) Mean DPG scores. (b) Paired differences (TSGate minus Gated attention) with pointwise source-bootstrap 95% CIs; the dashed line marks zero. Raw: original; Str: structured; Hdr: header only; Rep: semantic repeat; Pad: neutral padding; REC: full recaption.
Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resolutions beyond their training range. Existing training-free approaches mitigate this by modifying inference-time attention behavior, often through Rotary Position Embeddings (RoPE) extrapolation combined with attention scaling. However, these strategies apply a uniform and content-agnostic scaling across RoPE components with distinct frequency characteristics, inducing a trade-off between preserving global structure and recovering fine detail. We introduce SEGA, a training-free method that dynamically scales attention across RoPE components according to the latent's spatial-frequency structure at each denoising step. This adaptive scaling improves both structural coherence and fine-detail fidelity. Experiments show that SEGA consistently improves high-resolution synthesis across multiple target resolutions, outperforming state-of-the-art training-free baselines.
Diffusion Transformers (DiTs) are a dominant backbone for high-fidelity text-to-image generation due to strong scalability and alignment at high resolutions. However, quadratic self-attention over dense spatial tokens leads to high inference latency and limits deployment. We observe that denoising is spatially non-uniform with respect to aesthetic descriptors in the prompt. Regions associated with aesthetic tokens receive concentrated cross-attention and show larger temporal variation, while low-affinity regions evolve smoothly with redundant computation. Based on this insight, we propose AccelAes, a training-free framework that accelerates DiTs through aesthetics-aware spatio-temporal reduction while improving perceptual aesthetics. AccelAes builds AesMask, a one-shot aesthetic focus mask derived from prompt semantics and cross-attention signals. When localized computation is feasible, SkipSparse reallocates computation and guidance to masked regions. We further reduce temporal redundancy using a lightweight step-level prediction cache that periodically replaces full Transformer evaluations. Experiments on representative DiT families show consistent acceleration and improved aesthetics-oriented quality. On Lumina-Next, AccelAes achieves a 2.11× speedup and improves ImageReward by +11.9% over the dense baseline. Code is available at https://github.com/xuanhuayin/AccelAes.
Xuanhua Yin, Chuanzhi Xu, Haoxian Zhou +2
School of Computer Science, The University of Sydney, NSW 2006, Australia
Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.
Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3
School of Electrical Engineering, Iran University of Science & Technology, Tehran, IRAN · Independent researcher