We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target. Because we index these targets by a continuous depth coordinate, a trained model can loop far beyond its training depth with no retraining, early exits, or other modifications. In our experiments, these longer rollouts improve generation, so inference computation can grow without adding parameters. On ImageNet at 256x256, LiFT-L/2 achieves an FID 3.34 points lower than our dense DiT-XL/2 baseline while using approximately 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.
Figures & tables
Figure 1: Looping a smaller model beats larger dense DiTs at L/2 and XL/2 scale. FID against inference FLOPs per parameter for LiFT checkpoints at 50 integration steps (solid; colour encodes inference depth, marker area parameter count) and for dense DiTs as the number of steps varies (dotted). Only at B/2, the dense model remains stronger.
Figure 2: Loop in depth, flow in time. (a) The sampler advances through generative time, calling the network once per step. (b) Within each call, a shared core is applied K times, each time at its own depth coordinate sk , and a shared readout maps every state to a prediction uk . (c) In training, each uk is regressed onto its point uˉsk on the path from b=sg(u0) to u⋆ . (d) Training samples the interior coordinates at random; inference uses a uniform grid. Colours mark the same loop across panels; normalization and the first reinjection of h0 are omitted for clarity.
Model
LR
Ktrain
Lexectrain
N (M)
DiT-B/2
–
–
12
130.30
DiT-L/2
–
–
24
457.83
DiT-XL/2
–
–
28
674.82
LiFT-B/2-R1
1
8
12
56.69
LiFT-B/2-R4
4
2
12
88.58
LiFT-L/2-R5
5
4
24
175.79
Table 1: Same executed depth, far fewer parameters. Representative LiFT configurations and their dense references; LR is the shared-core block count and Lexectrain the executed training depth, including the LP=4 prelude blocks, and N the parameter count. Table 15 lists all configurations.
Figure 3: LiFT improves beyond its training depth, and its larger cores overtake dense DiTs. FID versus inference depth at 50 integration steps; rings mark training depth and dotted lines the dense DiTs. At L/2 and XL/2, the largest cores keep improving well past training depth and overtake the dense model of their scale, whereas single-block cores plateau. Training FLOPs match within each scale.
Model
Ktrain
FID at Ktrain
Best Kinf
Best FID
r
B/2 R4
2
37.95
16
33.87
8×
L/2 R10
2
20.30
8
10.95
4×
XL/2 R12
2
17.07
16
9.31
8×
XL/2 R8
3
16.95
32
9.50
10.67×
L/2 R5
4
21.42
32
16.21
8×
B/2 R1
8
45.70
16
45.58
2×
Table 2: The best inference depth is often far beyond the training depth. FID at training depth Ktrain and best FID over inference depth Kinf , at ratio r=Kinf/Ktrain ; Table 4 lists all checkpoints.
Figure 4: Looping more reaches lower FID than sampling more for L/2 and XL/2. FID versus inference cost: LiFT varies loops at 50 integration steps (solid), and dense DiTs vary the number of steps (dotted); rings mark training depth. Training FLOPs match within each model scale only.
Figure 5: Balancing inference loops and integration steps gives the best quality per budget. (a) Lowest measured FID within each inference budget for LiFT L/2 R10, which varies both, and for dense DiTs, which vary integration steps; labels give (T,Kinf) . (b) FID across inference loops and integration steps, interpolated between the measured dots; dashed lines connect allocations of equal cost, the white line marks the cost of dense L/2 at 50 steps (FID 16.94), and the pink path traces the best allocation as the budget grows. Dense L/2 matches LiFT’s training FLOPs.
Figure 6: LiFT needs fewer parameters for every quality target, and less compute for strict ones. Least inference compute (left) and fewest parameters (right) with which any evaluated LiFT or dense setting reaches each target FID; settings span scales with different training FLOPs.
Figure 7: More loops refine samples from the same noise. Selected paired samples from LiFT L/2 R10 at 50 integration steps and dense L/2 at varying numbers of steps; each row shares class and initial noise, and headers give inference TFLOPs per image. Compare the second column with the last: at eight inference loops, LiFT uses 30% less compute than dense L/2 at 250 steps and reaches a lower FID (10.95 versus 15.74).
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Kinf
T
TFLOPs/image
FID
sFID
IS
–
1
0.046
321.17
408.50
2.04
–
10
0.460
44.49
13.18
45.92
Dense B/2
–
25
1.150
32.50
8.92
53.42
–
50
2.300
30.10
7.84
54.64
–
100
4.600
29.22
7.39
55.11
–
250
11.501
28.82
7.19
54.86
Appendix
Table 3: Complete numerical results corresponding to the main figures. Each row reports FID, sFID, and IS from the same 50,000 generated images. R denotes the number of blocks in the shared core; a dash denotes a dense model.
Model
Ktrain
FID at Ktrain
Best Kinf
Best FID
r
B/2 R4
2
37.95
16
33.87
8×
L/2 R10
2
20.30
8
10.95
4×
XL/2 R12
2
17.07
16
9.31
8×
XL/2 R8
3
16.95
32
9.50
10.67×
B/2 R2
4
41.16
16
37.18
4×
L/2 R5
4
21.42
32
16.21
8×
Appendix
Table 4: Depth extrapolation for all evaluated checkpoints. All settings use 50 integration steps and 50,000 generated images. The best tested depth minimizes FID within each checkpoint’s measured sweep; r=Kinf/Ktrain is evaluated at that depth.
Figure 8: FID of every checkpoint in the main depth sweeps across inference depths, on a shared colour scale. The outlined cell in each row marks training depth, and grey cells indicate unevaluated settings.
Kinf
Model
Ktrain
1
2
3
4
5
6
8
10
12
16
20
24
32
B/2 R4
2
42.40
37.95
37.37
35.10
35.11 †
–
34.15
–
–
33.87
–
–
33.96 †
r
–
0.50
1.00
1.50
2.00
2.50
3.00
4.00
5.00
6.00
8.00
10.00
12.00
16.00
L/2 R10
2
31.83
20.30
14.64
12.25
11.70
–
10.95
–
–
10.99 †
–
–
11.57 †
r
–
0.50
1.00
1.50
2.00
2.50
3.00
4.00
5.00
6.00
8.00
10.00
12.00
16.00
XL/2 R12
2
27.71
17.07
12.74
10.69
10.25
–
9.55
–
–
9.31
–
–
9.71 †
Appendix
Table 5: FID of every checkpoint across inference depths, with the depth ratio r beneath each row. Bold values mark training depth, dashes denote unevaluated settings, and † marks an increase in FID relative to the preceding evaluated depth.
Figure 9: FID across integration steps and recurrent depth. Colour encodes FID on a logarithmic scale.
T
Kinf
TFLOPs/image
FID
sFID
IS
10
1
0.942
48.613
10.751
41.770
2
1.614
33.633
11.497
58.830
4
2.959
18.955
6.107
81.803
8
5.648
15.289
5.428
90.433
16
11.027
15.094
5.481
90.670
32
21.785
16.435
5.606
86.694
Appendix
Table 6: Complete joint-sweep operating points for LiFT L/2 R10, including additional depths 3 and 5 at 50 integration steps. All metrics in a row use the same images.
Budget (TFLOPs)
LiFT T
LiFT K
LiFT FID
Dense setting
Dense FID
0.046
–
Dense B/2, T=1
321.169
0.161
–
–
–
Dense L/2, T=1
321.150
0.460
–
Dense B/2, T=10
44.486
0.942
48.613
Dense B/2, T=10
44.486
1.150
1
48.613
Dense B/2, T=25
32.501
1.614
48.613
Dense L/2, T=10
29.591
Appendix
Table 7: FID-minimizing measured settings within each inference budget. The dense column selects among both evaluated dense backbones. Budgets are rounded for display; nearby distinct costs may share a printed value.
Figure 10: Training-grid randomization improves depth robustness. Both models use trajectory supervision with λ=0 and are evaluated on uniform grids ending at s=1 . Each point uses the same 50,000 initial noise samples and class labels, EMA weights, 50 Euler steps, and no classifier-free guidance. Hollow rings mark the training depth, K=4 .
Random s
Fixed s
Kinf
FID ↓
sFID ↓
IS ↑
FID ↓
sFID ↓
IS ↑
1
98.48
15.80
14.58
151.23
57.03
6.64
2
78.57
10.75
18.19
99.47
13.12
14.88
3
70.87
9.00
19.32
75.81
8.07
18.48
4
68.51
9.43
19.49
68.31
9.45
19.55
5
67.46
9.51
19.47
66.35
10.65
19.74
Appendix
Table 8: Random versus fixed training coordinates. Complete metrics for B/2 R2 with Ktrain=4 , λ=0 , and 100,000 updates. All settings use 50,000 images and 50 integration steps.
Figure 11: Prelude supervision controls the scale of the initial prediction. LiFT B/2 R2, trained with K=4 for 100,000 updates, varying only the auxiliary weight λ . Training panels show logged values faintly and an 11-point moving median prominently, omitting the first 2,000 updates; validation curves show unsmoothed full-validation EMA velocity MSE. The dashed line marks the approximate target RMS.
λ
Anchor RMS
Prelude MSE
Gradient norm
Validation MSE
0
3.3875
7.9626
0.0978
0.777666
0.01
0.9238
0.8316
0.0568
0.775813
0.1
0.9311
0.8100
0.0661
0.776069
1
0.9380
0.7992
0.1525
0.777948
Appendix
Table 9: Prelude-loss training diagnostics. Anchor RMS, prelude MSE, and gradient norm are medians of logged values over updates 90,000–100,000. Validation MSE uses EMA weights at update 100,000.
Figure 12: Generation quality across prelude-loss weights. FID, sFID, and IS for the same four 100,000-update checkpoints used in Figure 11 . Each point uses 50,000 generated images with 50 integration steps, with common initial noise and class labels. Hollow rings mark the training depth, K=4 .
λ
Metric
K=1
2
3
4
5
8
16
32
0
FID
98.48
78.57
70.87
68.51
67.46
66.97
67.71
68.38
sFID
15.80
10.75
9.00
9.43
9.51
9.43
9.26
9.26
IS
14.58
18.19
19.32
19.49
19.47
19.31
19.05
18.92
0.01
FID
93.90
75.69
69.22
66.39
65.24
64.81
65.46
66.01
sFID
16.55
11.86
10.21
9.23
8.66
8.35
8.39
8.47
IS
15.47
18.87
20.00
20.29
20.33
20.27
20.16
20.07
Appendix
Table 10: Complete prelude-loss sweep. All runs use the same B/2 R2 architecture, Ktrain=4 , and 100,000 updates. Each setting uses 50,000 images, EMA weights, 50 Euler steps, and no classifier-free guidance.
Figure 13: Prediction trajectories of LiFT L/2 R10, trained with two loops. (a) Projected progress toward the sampled target follows the assigned depth coordinate. (b) Projection onto each rollout’s own endpoint chord is close to linear, and (c) the orthogonal residual remains similar across depths. Curves average five generative times within each of 512 images; shaded bands show approximate 95% intervals across images. Chord endpoints agree by construction.
Kinf
∣βk−sk∣
∥rk∥2/d
100∥rk∥2/∥uK−b∥2
100∥rk∥2/∥uK∥2
2
0.0006
0.0556
0.56
6.07
4
0.0030
0.0710
0.64
7.72
8
0.0040
0.0764
0.68
8.29
16
0.0049
0.0782
0.70
8.48
32
0.0051
0.0800
0.72
8.68
Appendix
Table 11: Intermediate deviations from each rollout’s own endpoint chord. We average over interior readouts, then generative times and images; the endpoints are excluded because they lie on the chord by construction. The last two columns use different normalization scales.
Grid
Mean (%)
Image-mean p95
Worst-time p95
Cosine
Target MSE
Uniform K=1
11.01±0.20
14.88
26.86
0.99198
0.7765
Uniform K=2
5.12±0.12
7.63
16.59
0.99805
0.7669
Uniform K=3
2.61±0.07
4.08
8.92
0.99948
0.7668
Uniform K=4
0.00±0.00
0.00
0.00
1.00000
0.7670
Uniform K=5
1.70±0.05
2.80
6.64
0.99977
0.7675
Uniform K=8
2.65±0.10
4.88
13.98
0.99927
0.7684
Appendix
Table 12: Paired endpoint agreement with the uniform K=4 prediction. Differences are percentages of the reference prediction norm; means have approximate 95% intervals across 512 images. The image-mean p95 averages the five times within each image before taking the percentile, whereas the worst-time p95 first takes their maximum. All nonuniform grids have K=4 . MSE is measured against u⋆ .
Figure 14: Sensitivity of the final velocity to the computational grid, relative to uniform K=4 at the identical noisy input. The left panel shows the mean relative distance and the right shows its 95th percentile across 512 images at each generative time. Solid curves change coordinates while holding K=4 fixed; the dashed curve also increases the number of recurrent applications to 32.
t
BF16–FP32 RMSE
Front/ ϵ
Back/ ϵ
Random 1/ ϵ
K=32 / ϵ
0.05
0.00663
1.2
0.9
0.8
1.2
0.25
0.00654
3.5
4.3
2.4
5.9
0.50
0.00717
3.5
3.0
2.7
7.5
0.75
0.00937
2.4
2.5
1.9
4.9
0.95
0.01500
1.5
1.4
1.1
2.8
Appendix
Table 13: Numerical precision check on the same 16 examples at each time. The second column gives pooled BF16–FP32 endpoint RMSE for uniform K=4 . The remaining columns divide each grid’s pooled RMSE relative to that BF16 reference by the numerical baseline, using the identical examples.
Backbone
Width
Attention heads
Dense depth
B/2
768
12
12
L/2
1024
16
24
XL/2
1152
16
28
Appendix
Table 14: Backbone dimensions shared by the dense and recurrent models. Dense depth counts transformer blocks.
Model
LR
Ktrain
Lunique
Lexectrain
N (M)
Ctrain (EFLOPs)
Dense references
DiT-B/2
–
–
12
12
130.30
17.666
DiT-L/2
–
–
24
24
457.83
61.973
DiT-XL/2
–
–
28
28
674.82
91.098
LiFT configurations
LiFT-B/2-R1
1
8
5
12
56.69
17.697
Appendix
Table 15: Complete model configurations and training resources for the checkpoints in Section A.1 , with R denoting the shared-core block count. All LiFT models use four prelude blocks and no transformer coda, giving Lunique=4+LR unique blocks and Lexectrain=4+KtrainLR executed blocks per training forward pass. Executed depth matches the dense reference within each backbone scale, with its shared value centered across each LiFT group. Parameter counts N are obtained from the implementation and include conditioning and readout parameters, whereas block counts omit this work. The analytic forward-plus-backward estimate Ctrain accounts for it using the protocol in Section D.2 and the accounting in Section D.5 , and is reported in EFLOPs ( 1018 FLOPs).
Setting
Value
Optimizer
AdamW
Training updates
500,000
Global batch size
256
Learning rate
10−4
Schedule
500 warmup updates, then constant
Adam coefficients
(0.9,0.999)
Appendix
Table 16: Shared settings for the main training protocol.
Backbone
LR
Ktrain
Dense (EFLOPs)
LiFT (EFLOPs)
Extra cost
B/2
2
4
17.6655
17.6812
0.0889%
L/2
10
2
61.9726
61.9842
0.0188%
XL/2
4
6
91.0976
91.1391
0.0455%
Appendix
Table 17: Analytic training compute for representative matched-depth configurations. Each model uses 500,000 updates and a global batch size of 256; one EFLOP is 1018 operations. LiFT uses LP=4 and LD=0 . The totals include the forward and approximate backward cost of conditioning and all trajectory readouts. Overheads are computed before rounding the totals.
Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.
Yong Xien Chng, Tianyi Chen, Wenwen Tong +7
SenseTime Research · LeapLab, Tsinghua University · Nanyang Tehnological Univesity
To address the high sampling cost of Diffusion Transformers (DiTs), feature caching offers a training-free acceleration method. However, existing methods rely on hand-crafted forecasting formulas that fail under aggressive skipping. We propose L2P (Learnable Linear Predictor), a simple data-driven caching framework that replaces fixed coefficients with learnable per-timestep weights. Rapidly trained in ~20 seconds on a single GPU, L2P accurately reconstructs current features from past trajectories. L2P significantly outperforms existing baselines: it achieves a 4.55x FLOPs reduction and 4.15x latency speedup on FLUX.1-dev, and maintains high visual fidelity under up to 7.18x acceleration on Qwen-Image models, where prior methods show noticeable quality degradation. Our results show learning linear predictors is highly effective for efficient DiT inference. Code is available at https://github.com/Aredstone/L2P-Cache.
Zhirong Shen, Rui Huang, Jiacheng Liu +6
Shanghai Jiao Tong University · University of Electronic Science and Technology of China · Shandong University +2
Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, FlashLoop delivers lossless accuracy while achieving up to 1.64× end-to-end speedup and up to 6× KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
Wanqi Yang, Shiwei Liu
ELLIS Institute Tübingen Max Planck Institute for Intelligent Systems Tübingen AI Center