Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.
Figures & tables
Figure 1 : Layer use and tokenizer training in RAEv2, DRoRAE, and HiRAE-24.
Figure 2 : HiRAE-24 architecture and training. (a) Layer-wise experts and signed routing combine all 24 encoder layers under depth-dependent residual controls. (b) Tokenizer training jointly updates fusion and decoder while freezing the encoder. (c) Generator training freezes the tokenizer.
Figure 3 : Matched reconstruction details. Each selected triplet shows the input, official RAEv2, and HiRAE-24. Red boxes mark corresponding regions, enlarged below.
Tokenizer
Configuration
rFID ↓
PSNR (dB) ↑
LPIPS ↓
SD-VAE [ 24 ]
f8, 4 channels
0.610
26.90
0.130
VA-VAE [ 24 ]
f16, 32 channels
0.280
27.96
0.096
REPA-E [ 12 ]
VA-VAE + E2E tuning
0.280
26.25
0.110
RPiAE [ 9 ]
Pivot + variational bridge
0.500
21.30
0.216
FAE [ 7 ]
32-channel feature AE
0.680
–
–
RAE [ 27 ]
DINOv2-B, last layer
0.570
18.80
0.256
Table 1 : ImageNet-256 reconstruction. The upper block follows each source’s evaluation protocol; dashes denote unreported values. The lower block combines 50K rFID with PSNR/LPIPS on our matched 5K subset (100 classes; AlexNet LPIPS). Bold marks the best result within the lower block.
Figure 4 : Guided samples from HiRAE-24. Twenty-four selected class-conditioned images form eight visual groups, each with one larger example and two related samples.
Pretraining
Finetuning
Model
GenEval ↑
DPG ↑
GenAI ↑
GenEval ↑
DPG ↑
GenAI ↑
FLUX-VAE [ 1 ]
49.73
78.86
63.65
84.13
83.48
68.75
RAEv2 [ 20 ]
56.42
81.22
67.02
84.86
84.90
71.69
HiRAE-7
58.77
81.60
67.13
85.17
85.05
72.02
HiRAE-24
60.94
82.73
68.08
87.70
86.35
72.66
Table 4 : Text-to-image generation before and after SFT. DPG and GenAI denote DPG-Bench and GenAI-Bench. Bold denotes the best available score within each stage.
Table 6 : Fusion methods in the RAEv2 framework. DRoRAE-style fusion uses global interpolation to mix the deepest-layer representation and aggregated expert output at a fixed 80:20 ratio.
Figure 5 : Latent structure and decoding sensitivity. (a) Independently fitted PHATE views of the same images; colors denote classes. Same-class 10NN uses the original feature space. (b) Within HiRAE-24, lines connect anchor–fusion effective ranks for 100 images; points and error bars show means and 95% bootstrap CIs. (c) Mean output LPIPS under 10% relative latent perturbations.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Layers
Norm cap
Dropout
Shallow
0–7
0.025
0.50
Middle
8–15
0.075
0.25
Deep
16–23
0.150
0.10
Appendix
Table 9 : Residual-control settings for HiRAE-24. Layer ranges use zero-based indices.
Setting
Stage 1
Stage 2
Epochs
16
80
Global batch size
128
1024
Gradient accumulation
1
2
EMA decay
0.9978
0.9995
Peak learning rate
2×10−4
2×10−4
Final learning rate
2×10−5
2×10−5
Appendix
Table 10 : Training settings for HiRAE-24.
Setting
Pretraining
SFT
Hardware
8× H800 80GB
8× H800 80GB
Precision
BF16 mixed precision
BF16 mixed precision
Microbatch per GPU
32
32
Gradient accumulation
4
4
Global batch size
1024
1024
Optimizer updates
100,000
2,850
Appendix
Table 11 : Text-to-image training settings. Both phases train the same generator architecture with frozen visual and text encoders.
System
Generator
Parameters (M)
DiT / SD-VAE
DiT-XL/2
675
SiT / SD-VAE
SiT-XL/2
675
REPA
SiT-XL/2
675
VA-VAE
LightningDiT-XL
675
REPA-E / E2E-VAE
SiT-XL/2 + REPA
675
RPiAE
LightningDiT
675
Appendix
Table 12 : Generator configurations for the main-table baselines. Parameter counts exclude tokenizers and separate guidance models.
System
Unguided gFID
Guided gFID
Guided setting
RAE
2.160
1.740
AG
DRoRAE
2.680
1.650
AG
REPA-E / E2E-VAE
3.460
1.670
CFG
RPiAE
2.250
1.510
AG ∗
FAE, timestep shift
2.080
1.700
CFG
DecQ, 8 queries
1.800
1.330
AG
Appendix
Table 13 : ImageNet-256 generation after 80 generator epochs. Reference systems retain their reported sampling settings; dashes denote unavailable results. ∗ RPiAE’s guidance label differs between the source’s text and table.
Method
Epochs
gFID
FDr6
SiT-XL/2
800
2.120
8.440
DDT-XL
800
1.260
5.700
SiT-XL/2 + REPA
800
1.420
5.450
LightningDiT
800
1.420
4.570
REG
800
1.540
4.640
REPA-E
800
1.120
3.040
Appendix
Table 14 : Guided generation measured in six representation spaces. Baselines are transcribed from RAEv2 Table 7 and compared with HiRAE-24.
Reference
Δ PSNR
95% interval
Δ LPIPS ×103
95% interval
RAEv2
+3.710
[3.681,3.742]
−31.039
[−31.327,−30.757]
HiRAE-7
+1.309
[1.292,1.327]
−6.389
[−6.523,−6.257]
Appendix
Table 15 : Paired reconstruction improvements. HiRAE-24 minus each reference on the same 5,000 images. Intervals are 95% within-class paired image-bootstrap intervals; negative LPIPS differences indicate improvement.
Tokenizer
Sobel MSE ↓
Laplacian MSE ↓
RAEv2
1.830
43.378
HiRAE-7
1.167
36.124
HiRAE-24
0.816
32.645
Appendix
Table 16 : Reconstruction of spatial detail. Mean derivative errors on the same matched 5K subset as Table 1 , multiplied by 103 . Lower is better.
Figure 6 : Texture-stratified reconstruction. Quartiles increase in original-image Sobel strength. Left: LPIPS reduction from RAEv2 to HiRAE-24. Right: LPIPS increase after removing HiRAE-24’s shallow group. Each quartile contains 1,250 images; error bars show 95% paired bootstrap intervals.
HiRAE-24
HiRAE-7
Encoder layers used
All 24
Selected 7
Expert parameters (M)
≈ 202
≈ 59
Reconstruction rFID ↓
0.209
0.217
PSNR (dB) ↑
26.377
25.068
LPIPS ↓
0.043
0.049
Guided gFID ↓
1.038
1.038
Appendix
Table 17 : Selected-layer extension of HiRAE. Expert parameters exclude the encoder, router, decoder, and generator. rFID uses the 50K reconstruction evaluation; PSNR/LPIPS use the matched 5K subset. Generation uses 50K samples from epoch-80 generators.
Metric
DRoRAE-style fusion
HiRAE-24
Reduction
20 generator epochs
gFID ↓
2.244
2.242
0.1%
FDr6↓
5.270
2.533
51.9%
80 generator epochs
gFID ↓
1.551
1.038
33.1%
FDr6↓
3.375
1.856
45.0%
Appendix
Table 18 : Full-depth composition across training budgets. Both configurations use all 24 layers and 24 experts. Reductions use the adapted DRoRAE-style fusion as the reference.
Groups
Reconstruction
Unguided
IG=1.78
rFID ↓
PSNR ↑
LPIPS ↓
gFID ↓
IS ↑
gFID ↓
IS ↑
2
3.293
22.417
0.1725
44.795
42.976
29.385
62.103
3
3.363
22.153
0.1754
44.055
43.216
28.354
63.796
4
3.457
22.065
0.1789
45.172
41.618
31.321
59.984
Appendix
Table 19 : Three depth groups give the best generation results. Matched budgets: 16,000 tokenizer updates and 5,004 generator updates. Reconstruction uses 5K images; each generation setting uses 10K samples.
Groups
Layers per group
Norm caps
Dropout probabilities
2
12/12
0.0625,0.1875
5/12,0.15
3
8/8/8
0.025,0.075,0.15
0.5,0.25,0.1
4
6/6/6/6
0.01875,0.04375,0.075,0.1125
0.5,1/3,0.2,0.1
Appendix
Table 20 : A shared depth prior across group counts. Entries follow increasing encoder depth.
Group
Norm cap
LPIPS increase
Paired interval
Shallow (0–7)
0.025
5.650
[5.563, 5.737]
Middle (8–15)
0.075
54.111
[53.600, 54.626]
Deep (16–23)
0.150
147.750
[146.789, 148.661]
Appendix
Table 21 : Each residual group contributes to reconstruction. We remove one group from the frozen HiRAE-24 tokenizer. Values report ΔLPIPS×103 with paired 95% intervals.
Group
High removal − energy control
High − low, equal norm
Paired interval
Shallow
0.770
0.215
[0.182, 0.248]
Middle
4.760
0.068
[ −0.013 , 0.144]
Deep
7.328
−2.053
[ −2.163 , −1.941 ]
Appendix
Table 22 : Depth groups differ in frequency contributions. Values report differences in reconstruction LPIPS, multiplied by 103 . The equal-norm comparison uses κ=1 ; positive values indicate greater damage from high-frequency removal.
Coordinates
Representation
High-frequency fraction
Same-class 10NN
Shared anchor
Deep anchor
0.123
0.827
Shared anchor
Fused representation
0.208
0.815
Raw
Deep anchor
0.107
0.796
Raw
Fused representation
0.166
0.795
Appendix
Table 23 : Spatial enrichment largely preserves class neighborhoods. Values compare HiRAE-24’s deep anchor and fused representation under shared anchor normalization and without additional standardization.
Figure 7 : Cross-system spatial statistics. Left: class-mean adjacent-patch cosine similarity. Right: high-frequency DCT energy fraction without DC. We standardize each representation per channel and remove spatial means for the cosine measure.
Tokenizer
Gaussian
Low frequency
High frequency
RAEv2
8.385
7.191
9.414
HiRAE-7
1.576
1.344
1.810
HiRAE-24
1.791
1.547
2.043
Appendix
Table 24 : HiRAE decodings change less under the tested perturbations. Output LPIPS ×103 at 10% relative latent perturbation norm.
Figure 8 : Prediction error varies with noise level. Full-branch clean-latent MSE on matched forward-noised images. HiRAE-24 improves over RAEv2 at the five interior levels; both endpoints have slightly higher error.
Coefficient log-SNR
RAEv2
HiRAE-7
HiRAE-24
−6
0.459
0.617
0.473
−4
0.296
0.399
0.292
−2
0.171
0.224
0.165
0
0.103
0.112
0.088
2
0.052
0.047
0.042
4
0.018
0.017
0.016
Appendix
Table 25 : Complete prediction-error grid. Per-coefficient clean-latent MSE, including DC, at all seven coefficient log-SNR levels.
Figure 9 : Additional reconstruction comparisons. Each triplet shows the input, official RAEv2 reconstruction, and HiRAE-24 reconstruction. The first two rows contain four selected cases; the remaining six rows contain twelve cases at predetermined validation indices.
Figure 10 : Unguided class-conditioned samples from HiRAE-24. The first eight stored samples are shown without quality filtering, using the epoch-80 EMA model with CFG=IG=1.
Figure 11 : Qualitative comparisons on five selected GenEval prompts. Each row shows the prompt and the outputs of FLUX-VAE, RAEv2, HiRAE-7, and HiRAE-24 after SFT. We display the complete original images. FLUX-VAE denotes the FLUX autoencoder paired with our trained DiT.
Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insights which simplify and improve RAE. First, we study a generalized formulation where the representation is defined as sum of the last k encoder layers rather than solely the final layer. This simple change greatly improves reconstruction without encoder finetuning or specialized data (e.g., text, faces). Second, we study the prevalent assumption that RAE (using pretrained representation as encoder) replaces representation alignment (REPA), which distills the same representation to intermediate layers instead. Through large-scale empirical analysis, we uncover a surprising finding: RAE and REPA exhibit complementary working mechanisms, allowing the same representation to be used as both encoder and target for intermediate diffusion layers. Finally, the original RAE struggles with classifier-free guidance (CFG) and requires training a second, weaker diffusion model for AutoGuidance (AG). We show that REPA itself can be viewed as x-prediction in RAE latent space. By simply re-parameterizing the output of the DiT model, it can provide guidance for "free". Overall, RAEv2 leads to more than 10x faster convergence over the original RAE, achieving a state-of-the-art gFID of 1.06 in just 80 epochs on ImageNet-256. On FDr6, RAEv2 achieves a state-of-the-art 2.17 at just 80 epochs compared to the previous best 3.26 (800 epochs) without any post-training. This motivates EPFID@k (epochs to reach unguided gFID < k) as a measure of training efficiency. RAEv2 attains an EPFID@2 of 35 epochs, versus 177 for the original RAE. We also validate our approach across diverse settings for text-to-image generation and navigation world models, showing consistent improvements. The code is available at https://raev2.github.io.
Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spaces for image generation. However, their reconstruction quality often remains suboptimal, largely because deep VFM representations do not preserve sufficient fine-grained visual detail. This limitation becomes even more severe after discretization, where missing low-level information is difficult to recover. In fact, we observe that shallow VFM features retain considerably richer local appearance and structural detail, which complements the high-level semantics carried by deep features used in existing RAEs. Motivated by this complementary property, we propose Ideal, an In-depth Alignment framework for discrete representation autoencoding. By jointly aligning quantized tokens with both shallow and deep VFM features, Ideal enables the resulting discrete visual tokens to preserve both visual fidelity and rich semantics. Extensive experiments demonstrate that Ideal yields superior reconstruction performance, achieving 0.61 rFID on ImageNet and outperforming the previous best method by 0.28. When used for autoregressive image generation, Ideal further produces a gFID of 1.89, establishing a new state of the art for autoregressive image generation.
Yitong Chen, Zijie Diao, Junke Wang +5
Institute of Trustworthy Embodied AI, Fudan University · Shanghai Innovation Institute · University of Maryland, College Park
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off: shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling preserves the full-layer latent mean in expectation while explicitly penalizing sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. Applying FuseReg to both stages also reduces unguided gFID by 29% on DiT-Base. The reconstruction and generation benefits also extend to other encoder families. FuseReg narrows the reconstruction-generation gap without additional training cost or architectural changes.
Hongyang Du, Yunfei Xie, Junjie Ye +13
USC PSI Lab · Brown University · Rice University +4