Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce \textit{REPresentation Injection} (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over 43.5×. Code will be available at https://jeneveuxpas.github.io/REPI
Figures & tables
Figure 1: REPI accelerates diffusion transformer training across model scales. Compared with REPA, REPI consistently achieves lower FID on SiT-B, SiT-L, and SiT-XL, and combining the two yields further gains. Notably, on SiT-XL, REPI + REPA at 200K training steps already outperforms REPA trained for 400K steps.
Injection target
FID ↓
Hidden state
212.5
Attention output
5.0
Q/K/V
5.2
K/V
7.1
K only
12.4
V only
4.6
Table 1: Oracle results (SiT-XL/2, 30K steps).
Figure 2: Overview of REPI. (a) Scaffold : projected keys and values from a frozen visual encoder temporarily replace the diffusion transformer’s native keys and values, while its native queries are preserved. (b) Internalization : the model resumes using its own keys and values, guided by an internalization objective that aligns them with the projected encoder representations. At inference time, the encoder and projection layers are discarded.
Table 4
Figure 3: Qualitative comparison across training iterations. Generated samples from SiT-XL/2 trained with iREPA (top) and REPI (bottom) over the first 400K iterations. Both models use the same initial noise and sampler, without classifier-free guidance.
Method
w/o CFG
w/ CFG
REPA
10.40
4.73
REPI
9.47
4.61
REPA + REPI
9.08
4.56
Table 4: FID comparison on text-to-image generation.
Figure 4: Generality of REPI across visual encoders. FID results of REPA, REPI, and REPA + REPI at 100K training steps on SiT-XL/2. REPI consistently outperforms REPA across all ten encoders. Notably, on encoders such as MAE-L and MoCoV3-L, where REPA underperforms vanilla SiT, REPI still achieves substantial gains.
Figure 5: Robustness of REPI across diffusion backbones and injection layers. (a–c) FID comparison on DiT-L/2, DiG-L/2, and SiT-XL/2 ( 512×512 ). REPI consistently outperforms REPA, and combining the two yields further gains. (d) REPI is robust to the choice of injection depth, with layer 8 achieving the best performance on SiT-XL/2 at 100K steps.
Table 6: Hyperparameters and model configurations for standalone REPI , shared across all ImageNet 256×256 and 512×512 experiments.
Method
Time (h)
FID ↓
IS ↑
REPA
4.63
19.40
67.4
REPI
4.78 (+3.2%)
14.51 ( − 25.2%)
82.7 (+22.7%)
REPA + REPI
4.81 (+3.9%)
11.78 ( − 39.3%)
96.8 (+43.6%)
Appendix
Table 7: Total training time for 100K steps on SiT-XL/2 using four NVIDIA H200 GPUs. Percentages in parentheses denote the relative change with respect to REPA.
Figure 8: Illustration of injection targets within a DiT block.
Layer
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
2
12.62
5.66
91.56
0.70
0.61
4
11.78
5.58
96.77
0.70
0.61
6
12.79
5.64
91.55
0.69
0.61
8
14.70
5.79
83.75
0.68
0.61
10
14.61
5.78
82.81
0.68
0.61
12
14.02
5.76
84.30
0.68
0.61
Appendix
Table 8: Ablation results on the REPI injection layer when combined with REPA. The REPA alignment layer is fixed at its default, layer 8. All results are reported on SiT-XL/2 at 100K training steps on ImageNet 256×256 without classifier-free guidance.
λ
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
0.125
12.13
5.80
95.17
0.70
0.61
0.25
11.78
5.58
96.77
0.70
0.61
0.5
12.02
6.37
95.13
0.70
0.60
1.0
12.68
7.05
93.18
0.69
0.60
Appendix
Table 9: Ablation results on the internalization loss weight when combined with REPA. All results are reported on SiT-XL/2 at 100K training steps on ImageNet 256×256 without classifier-free guidance.
Injection target
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
Attention output
16.34
5.89
78.62
0.67
0.60
Q/K/V
17.40
5.91
71.60
0.67
0.60
K only
15.81
5.89
77.57
0.67
0.60
V only
15.53
5.52
79.31
0.68
0.60
K/V
14.51
5.57
82.71
0.68
0.60
Appendix
Table 10: Ablation results on the injection target. All results are reported on SiT-XL/2 at 100K training steps on ImageNet 256×256 without classifier-free guidance.
Model
#Params
Iter.
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
SiT-B/2 ( Ma et al., 2024 )
130M
400K
33.0
6.50
43.7
0.53
0.63
+ REPA
130M
100K
49.5
7.00
27.5
0.46
0.59
+ REPA
130M
200K
33.2
6.68
43.7
0.54
0.63
+ REPA
130M
400K
24.4
6.40
59.9
0.59
0.65
+ REPI
130M
100K
35.9
6.65
40.3
0.54
0.62
+ REPI
130M
200K
25.0
6.74
59.6
0.59
0.65
Appendix
Table 11: Detailed quantitative results across different SiT and DiT models. All results are obtained on ImageNet 256×256 without classifier-free guidance.
Figure 9: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “great white shark” (2).
Figure 10: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “loggerhead sea turtle” (33).
Figure 11: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “macaw” (88).
Figure 12: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “Blenheim Spaniel” (156).
Figure 13: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “Border collie” (232).
Figure 14: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “Arctic wolf” (270).
Figure 15: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “castle” (483).
Figure 16: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “convertible” (511).
Figure 17: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “ice cream” (928).
Figure 18: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “cheeseburger” (933).
Figure 19: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “buckeye” (990).
Figure 20: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “cliff” (972).
Figure 21: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “coral reef” (973).
Figure 22: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “lakeshore” (975).
Figure 23: Generated samples from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ), conditioned on class “volcano” (980).
Figure 24: Generated samples ( 512×512 ) from REPA + REPI on SiT-XL/2, trained for 400K steps, using classifier-free guidance ( w=4.0 ).