Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.
Figures & tables
Figure 1 : Looping is a parameter-efficient way to scale text-to-image generation models. (a) Increasing loop depth greatly improves text-to-image reasoning performance. (b) These gains are achieved with far fewer parameters and lower inference cost. (c) Self-correction emerges across loops.
Figure 2 : Overview of Looped-DiT. Shared middle blocks are repeated N times within each denoising step. Self-Modulating Attention (left) regulates attention updates with token-dependent, headwise gates. Deep Supervision (right) decodes each intermediate loop output through the shared post-loop blocks, supervising all predictions against the same clean-image target.
Figure 3 : Limitations of naïve looping in MMDiT. (a) Increasing inference loop depth beyond that used in training initially improves performance, then saturates and degrades. (b) This degradation coincides with a steady loss of linearly decodable spatial information under a ridge-regression probe.
Model
Params
GenEval
DPG
PRISM
CoRe
Spatial
TIIF-Short
Average
Non-CoT models
E-MMDiT [ 43 ]
0.30B
69.1
80.4
51.6
30.7
45.7
63.4
56.8
SANA-0.6B [ 55 ]
0.59B
65.8
83.3
56.4
39.8
47.4
69.4
60.4
DeCo-XXL/16 [ 34 ]
1.1B
82.0
82.0
52.6
34.8
49.8
70.4
61.9
URSA-0.6B [ 8 ]
0.86B
64.1
85.6
59.6
43.4
52.0
67.6
62.1
Table 1 : Comparison with state-of-the-art text-to-image models across six benchmarks.
Figure 4 : Performance-efficiency trade-offs across state-of-the-art text-to-image models. Average Score denotes the mean performance across 6 benchmarks. Blue circles denote Looped-DiT B/16 with varying loop depth and denoising steps.
Figure 5 : Qualitative results of Looped-DiT B/16. Across successive loops, the model resolves spatial constraints and corrects inconsistencies.
Model
Train GFLOPs
Inference GFLOPs
Params (M)
Effective depth
Hidden dim
DeepSup
Looping
Avg. ↑
MiniT2I B/32
441
146
260
17
768
✗
✗
55.2
Deeper MiniT2I B/32
809
267
473
32
768
✗
✗
57.7 (+2.6)
Wider MiniT2I B/32
805
268
489
17
1056
✗
✗
58.6 (+3.5)
Deeper MiniT2I B/32 w. DeepSup
1,246
267
473
32
768
✓
✗
58.1 (+3.0)
Looped-DiT B/32 (Ours)
1,246
267
260
32
768
✓
✓
59.1 (+4.0)
Table 2 : Benefits of looping under matched parameter and different compute budgets.
Figure 6 : Loop iterations versus additional denoising steps. We compare 4-loop inference with single-pass inference using the same looped checkpoint or the non-looped model. Both single-pass settings use additional denoising steps to match the per-image inference FLOPs of 4-loop inference.
Capability
CoT
Looping
Both
Benchmark avg.
+2.1
+4.0
+5.5
Constraint resolution (Looping-favored)
PRISM / Long Text
+3.1
+9.1
+9.4
CoRe / Multi-Relation
+1.2
+8.3
+8.4
CoRe / Procedural
+4.7
+9.4
+9.6
Table 3 : Complementary gains from looping and textual CoT over the non-looped model with original prompts. Both combines looping with CoT rewriting.
Looping
DeepSup
Attention
Benchmarks ↑
DPG
PRISM
CoRe
Spatial
Avg.
✗
✗
✗
82.0
51.2
39.2
48.3
55.2
✓
✗
✗
84.4
51.1
39.3
49.9
56.2
✓
Exponential
✗
84.3
52.1
41.0
51.0
57.1
✓
Final + Mean
✗
84.4
53.9
40.6
50.9
57.5
✓
✗
Gated
84.0
53.0
39.5
49.7
56.6
Table 4 : Component ablations. Avg. is the mean over the four reported benchmarks. DeepSup denotes deep supervision; ✗ indicates final-loop only supervision. The loop weights (w1,w2,w3,w4) are (\nicefrac18,\nicefrac14,\nicefrac12,1) for Exponential and (\nicefrac13,\nicefrac13,\nicefrac13,1) for Final + Mean. Gated and XSA are self-modulating attention variants.
Figure 7 : Deep supervision improves the standard 4-loop result, stabilizes early exits, and maintains performance beyond the training loop count. The model trains at 4 loops and infers at 1–8 loops.
Figure 8 : Adaptive looping results. (a) By choosing different loop counts for each case, adaptive looping produces correct images with lower average loop loop counts compared to fixed looping. (b) Adaptive looping unlocks compute–performance trade-off without retraining, and also reaches better performance with fewer per-image loops. The numbers denote average per-image loops.
Looping
Effective depth
Deep Sup
Attention
Avg. score ↑
ΔXSA
Non-looped, parameter-matched
✗
17
✗
✗
55.2
✗
17
✗
XSA
55.1
−0.1
Non-looped, compute-matched
Table 5 : Gains from XSA with and without looping. Non-looped depths 17 and 32 match the looped model’s parameter count and forward FLOPs, respectively. ΔXSA is the point gain over the paired variant without XSA.
Figure 9 : Self-modulating attention reduces excessive writing. (a) Ridge-regression R2 for decoding token positions from post-loop activations. (b) Attention-update norm relative to the residual-stream norm. (c) An example of XSA reducing redundant attention updates, and thus preventing an error that would be introduced by looping without XSA at the same denoising step.
Configuration
Looped-DiT B/32
Looped-DiT B/16
Image resolution
512×512
512×512
Patch size
32
16
Image tokens
256
1024
Joint sequence length
512
1280
Hidden dimension
768
768
Unique MMDiT blocks
17
17
Table 6 : Model configurations. Looped-DiT B/32 and Looped-DiT B/16 share the same Transformer backbone and differ mainly in patch size and the resulting image-token sequence length.
Looped-DiT B/32
Looped-DiT B/16
Hyperparameter
Pretrain
Fine-tune
Pretrain
Fine-tune
Training steps
250K
40K
500K
80K
Global batch size
1024
Nodes × GPUs
2×8
2×8
4×8
4×8
Optimizer
AdamW
Adam (β1,β2)
(0.9,0.95)
Table 7 : Training hyperparameters. Hyperparameters are shared across model scales and training stages unless otherwise noted.
Training
Inference
Configuration
Params
GFLOPs
GFLOPs
Compute-matched (Deeper)
473M (1.82×)
809 (1.83×)
267 (1.83×)
Compute-matched (Wider)
489M (1.88×)
805 (1.83×)
268 (1.84×)
Parameter-matched (MiniT2I-B/32)
260M (1.00×)
441 (1.00×)
146 (1.00×)
+ Looping
260M (1.00×)
809 (1.83×)
267 (1.83×)
+ XSA
260M (1.00×)
809 (1.83×)
267 (1.83×)
Table 8 : Training and inference cost. Training GFLOPs are measured per sample for one full training step, including forward and backward computation and, when applicable, the Deep Supervision exits. Inference GFLOPs are measured per denoising forward pass. Parenthesized values are relative to the parameter-matched MiniT2I-B/32 baseline.