Organizations: Central South University · The Hong Kong University of Science and Technology · Tsinghua University · The Chinese University of Hong Kong, Shenzhen · SenseTime Research · Shandong University · Imperial College London · University of Oxford · Dell Technologies · University of International Relations
Representation alignment speeds up diffusion transformer training by pulling an intermediate block of the model (student) toward features of a frozen pretrained encoder (teacher). Which teacher layer to align, and for how long, is still set by convention, and each alternative costs a training run. We find that alignment helps where the student cannot linearly recover the teacher's features, not where it already resembles them. Since a deep teacher layer is largely predictable from the one below, we isolate what each layer adds, its increment, and measure how much of it an unaligned student recovers. The student fills the teacher's hierarchy from the bottom up and stalls near the top, which we call hierarchy filling: even after 400K steps it recovers almost none of the deepest. The recoverability gap is the unrecovered share of an increment, read from one unaligned checkpoint. In short runs that each align one teacher layer at one block, the gap nearly reproduces their ranking by FID improvement, and CKA, a measure of feature similarity, largely reverses it. Representation Alignment and Recoverability Estimation (RARE) picks the teacher layer with the largest gap before training. During training, it tracks each token's remaining distance to that layer, the online counterpart of the gap, weights tokens by it, and phases out the loss once the average distance stops falling. With SiT-B/2 on ImageNet 256×256, RARE reaches an FID of 18.02 without guidance and 4.46 with it, ahead of seven alignment baselines including REPA, iREPA and HASTE. It also trains in 14% fewer GPU-hours than iREPA. Its FID stays below iREPA's across model scales, teachers, datasets and backbones.
Figures & tables
Figure 1: Alignment placed by convention and by measurement. (a) REPA alignment (REPAlign) pulls a fixed student block toward the teacher’s deepest layer for the whole run. (b) RARE first reads from a frozen checkpoint how much of each teacher layer the student still lacks, then applies sparse alignment (SPAlign): it aligns only the layer with the largest gap, puts more weight on tokens that are still far from the target, and switches the loss off once this distance stops falling, so alignment covers one layer and only the early part of training. (c) One class and noise draw at four training steps, guidance scale 4.0.
Figure 2: Hierarchy filling. The vertical axis is the held-out explained variance R2 of Eq. 2 , the share of a teacher target that a linear readout reconstructs from the frozen student, maximized over blocks and noise levels on the twelve-layer map. Higher means the student already encodes more of that target and leaves less for alignment to add. The grey band lies below τ=0.02 , the estimation noise of a held-out R2 . (a) Increments of four teacher layers over training. (b) Whole whitened layer, increment and clean-latent availability at 400K. (c) Three teachers at 100K; the shaded region spans SiT-B/2, SiT-L/2, SiT-XL/2 and Places365 under DINOv2-B.
Criterion
What It Measures
Spearman ρ with Benefit
Stable
10K
30K
50K
Recoverability gap N (Ours)
Unrecovered share of the target
+0.86
+0.90
+0.92
✓
Similarity to the Target
Readout R2
Recovered share of the target
−0.85
−0.87
−0.89
✓
CKA
Similarity to the target
−0.81
−0.84
−0.85
✓
Gram similarity
Match of token-to-token relations
+0.37
−0.37
−0.27
✗
Table 1: Placement criteria scored against realized alignment benefit. Spearman rank ρ between each criterion, read from the unaligned run at 10K, 30K and 50K steps, and the gFID benefit of the fourteen branches of Section 4.2 ; tied values share their mean rank. Stable marks a criterion whose percentile-bootstrap 95% interval excludes zero at all three checkpoints, in either direction.
Figure 3: The gap predicts alignment benefit; similarity predicts it in reverse. (a, b) Benefit of the fourteen branches against the gap and against CKA, on a logarithmic axis, on the 50K map, colored by teacher layer, with Spearman ρ and its 95% interval; the grey band marks benefits within 1.5 gFID of zero. (c) Benefit along the teacher axis at blocks 6 and 10.
Method
Alignment
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
CMMD ↓
Target
Block
Without Guidance
SiT ( Ma et al., 2024 ) (ECCV’24)
—
—
35.15
6.60
41.94
0.526
0.633
1.279
REPA ( Yu et al., 2025a ) (ICLR’25)
L12
4
22.39
6.62
65.62
0.594
0.648
1.067
iREPA ( Singh et al., 2026a ) (ICLR’26)
L12
4
20.40
6.90
72.61
0.602
0.650
1.054
HASTE ( Wang et al., 2026 ) (NeurIPS’25)
L12 + attn.
8
18.99
6.52
74.13
0.624
0.638
0.984
Table 2: ImageNet-256 with SiT-B/2 at 400K steps. L12 is DINOv2-B’s last layer, aligned raw (768 dimensions) unless marked whitened (128 dimensions); attn. is HASTE’s attention-map term.
Figure 4: Training dynamics. (a, b) FID-10K without guidance on a logarithmic axis against steps, mapped by (s/100K)1.2 to spread the early checkpoints, and against wall-clock hours; markers are measured checkpoints. (c) Gap at the selected cell (layer 12, block 6; twelve-layer map) at 200K and 400K, including the run that aligns this cell until 250K (Table 3 , row 4).
#
Target
Block
wk
Schedule
FID-50K ↓
no CFG
CFG
1
Whitened
4
✗
Always on
20.40
5.48
2
Whitened
6
✗
Always on
20.86
5.85
3
Increment
6
✗
Always on
21.79
6.42
4
Whitened
6
✗
Stop at 250K
19.32
4.90
5
Whitened
4
✗
Stop at 250K
18.83
4.77
Table 3: Ablation of RARE’s decisions at 400K steps. All rows align layer 12, whitened or as its increment, with RARE’s projector and loss weight. wk : token weights (Eq. 6 ); Measured: release of Eq. 7 ; Ramp: hand-set. Row 7 is RARE.
Setting
FID ↓
SiT
iREPA
RARE
(a) Scale, DINOv2-B, 100K Steps
SiT-B/2
62.59
37.70
36.66
SiT-L/2
47.19
19.94
18.93
SiT-XL/2
42.45
16.52
16.07
(b) Teacher, SiT-B/2, 100K Steps
Table 4: Transfer. Unguided FID-50K except the CFG row; RARE aligns layer 9 on Places365. Bold: better of iREPA and RARE.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Whole teacher layers are near-copies of one another; their increments are not. Top three principal components of each feature map as RGB, fitted jointly in a shared rank-128 basis. Left: the whole whitened layer separates object from background at every depth and changes mainly in hue. Right: the increments carry different content at each depth, from edges and texture at layer 3 to coarse, low-frequency regions at layer 12.
Figure 6: The recoverability gap map and its two axes. (a) Gap of the unaligned SiT-B/2 at 50K on all twelve teacher layers; hatched layers fall below the readout floor, and the box marks the cell the four-layer map selects. (b) Along the block axis, the gap of layer 12, the readout of the shallowest increment, and the benefit of aligning layer 12 at blocks 2, 6 and 10. (c) Spearman ρ of the nine single criteria at 50K with bootstrap 95% intervals, and at 10K and 30K.
Figure 7: Alignment closes the gap where it is applied. Gap closed relative to the unaligned run at 200K steps. (a) Along the teacher layers at block 6, for the run that aligns layer 12 at block 6 and for REPA, iREPA and HASTE. (b) Over the whole map for the first run; the box marks the selected cell.
Unaligned
Aligned
SiT
REPA
iREPA
HASTE
RARE
Parameters (M) ↓
130.3
137.6
135.6
137.6
131.2
Peak Memory (GiB/GPU) ↓
6.22
6.94
6.79
8.39
8.17
Step Time (s) ↓
0.168
0.218
0.218
0.212
0.188
Training Cost (GPU-h to 400K ) ↓
148.9
193.5
193.6
188.0
167.1
Inference Overhead
0
0
0
0
0
Appendix
Table 5: Training and inference cost of the runs of Table 2 . Every method discards its projector after training, so inference costs the same as for the unaligned model. Marks rank the aligned methods only.
Figure 8: The map’s cell comes within 1.37 gFID of the sweep’s best at a fraction of its cost. Realized gFID benefit of the cell each procedure picks against the GPU-hours spent picking it. Diagnosing once is 3681× cheaper than the sweep on selection alone, and 19.9× cheaper once the training of the selected cell is counted on both sides.
Setting
Method
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
CMMD ↓
Δ FID ↑
(a) Model Scale, DINOv2-B Teacher, 100K Steps
SiT-B/2
SiT
62.59
7.15
20.78
0.395
0.567
1.703
—
iREPA
37.70
6.95
40.19
0.514
0.618
1.408
24.89
RARE
36.66
7.32
40.04
0.531
0.616
1.359
25.93
SiT-L/2
SiT
47.19
6.30
27.65
0.482
0.586
1.355
—
iREPA
19.94
5.71
67.80
0.630
0.612
0.933
27.25
Appendix
Table 6: Transfer, all metrics. (a, b) are unguided; Δ FID is the improvement over the unaligned SiT of the same setting. RARE aligns layer 9 on Places365. The better of iREPA and RARE is in bold.
Setting
Whitened Layer
Increment
Layer
Block
FID ↓
Layer
Block
FID ↓
SiT-B/2, ImageNet
Layer 12
4
36.66
Layer 12
6
39.21
SiT-L/2, ImageNet
Layer 12
8
18.93
Layer 9
4
24.26
SiT-XL/2, ImageNet
Layer 12
8
16.07
Layer 9
5
20.15
SiT-B/2, Places365
Layer 9
4
7.33
Layer 9
2
9.19
Appendix
Table 7: Whitened layer against increment target. FID-50K without guidance, 100K steps on ImageNet and 400K on Places365. Each increment arm uses the cell its own map selects, with the same projector and loss weight. The better target is in bold.
Method
Alignment
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
CMMD ↓
Target
Block
Without Guidance
DiT
—
—
40.50
6.26
35.05
0.502
0.631
1.341
REPA
L12
4
30.05
6.35
49.86
0.556
0.650
1.203
iREPA
L12
4
24.16
6.36
63.68
0.586
0.643
—
HASTE
L12 + attn.
8
21.76
6.26
65.55
0.612
0.629
1.030
Appendix
Table 8: DiT-B/2 with ϵ -prediction and a 250-step DDPM sampler, 400K steps. Notation as in Table 2 .
Method
Alignment
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
CMMD ↓
Target
Block
Without Guidance
MM-DiT
—
—
56.19
7.45
25.88
0.419
0.593
1.688
REPA
L12
4
42.79
7.04
35.11
0.481
0.620
1.467
iREPA
L12
4
35.34
6.79
44.17
0.518
0.630
1.378
HASTE
L12 + attn.
8
34.28
7.38
45.10
0.537
0.628
1.332
Appendix
Table 9: MM-DiT-B/2 with flow matching, 100K steps. Notation as in Table 2 . A tie for best leaves no second best.
Figure 9: ImageNet-256 samples at 400K steps. Each column shares one class, initial noise and sampler noise across the rows, so differences within a column come from the model; guidance 4.0.
Figure 10: REPA and RARE over training. Three classes at 50K, 100K and 400K steps; guidance 4.0.
Figure 11: Places365-Standard at 400K steps. SiT, iREPA and RARE on shared columns; guidance 4.0.
Figure 12: One column through the sampler. Decoded prediction of the clean latent at seven points of the 250-step SDE; guidance 4.0. The rows share the noise and the first snapshot and separate at the second, where the three aligned rows already resolve the eye and the beak.