HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
Organizations: Peking University · Agibot Research · Tsinghua University · IGDL
Abstract
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.
Figures & tables
| Tokenizer | Configuration | rFID | PSNR (dB) | LPIPS |
|---|---|---|---|---|
| SD-VAE [ 24 ] | f8, 4 channels | 0.610 | 26.90 | 0.130 |
| VA-VAE [ 24 ] | f16, 32 channels | 0.280 | 27.96 | 0.096 |
| REPA-E [ 12 ] | VA-VAE + E2E tuning | 0.280 | 26.25 | 0.110 |
| RPiAE [ 9 ] | Pivot + variational bridge | 0.500 | 21.30 | 0.216 |
| FAE [ 7 ] | 32-channel feature AE | 0.680 | – | – |
| RAE [ 27 ] | DINOv2-B, last layer | 0.570 | 18.80 | 0.256 |
| System | Epochs | Guide | gFID | IS | |
|---|---|---|---|---|---|
| DiT-XL/2 [ 17 ] | 1400 | CFG | 2.270 | 278.200 | — |
| SiT-XL/2 [ 15 ] | 1400 | CFG | 2.060 | 270.300 | — |
| REPA [ 25 ] | 800 | CFG-int. | 1.420 | 305.700 | — |
| VA-VAE [ 24 ] | 800 | CFG | 1.350 | 295.300 | — |
| REPA-E [ 12 ] | 800 | CFG | 1.120 | 302.900 | — |
| RPiAE [ 9 ] | 80 | AG | 1.510 | 225.900 | — |
| Tokenizer | gFID | IS | |
|---|---|---|---|
| RAEv2 | 1.650 | 228.000 | 3.950 |
| RAEv2 K=23 | 3.010 | 206.000 | — |
| HiRAE-24 | 2.129 | 210.339 | 4.660 |
| Pretraining | Finetuning | |||||
|---|---|---|---|---|---|---|
| Model | GenEval | DPG | GenAI | GenEval | DPG | GenAI |
| FLUX-VAE [ 1 ] | 49.73 | 78.86 | 63.65 | 84.13 | 83.48 | 68.75 |
| RAEv2 [ 20 ] | 56.42 | 81.22 | 67.02 | 84.86 | 84.90 | 71.69 |
| HiRAE-7 | 58.77 | 81.60 | 67.13 | 85.17 | 85.05 | 72.02 |
| HiRAE-24 | 60.94 | 82.73 | 68.08 | 87.70 | 86.35 | 72.66 |
| Model | Raw-layer experts | Residual regularization | rFID | 20 epochs | 80 epochs | ||
|---|---|---|---|---|---|---|---|
| gFID | gFID | ||||||
| HiRAE | 0.209 | 2.242 | 2.533 | 1.038 | 1.856 | ||
| HiRAE | 0.230 | 2.410 | 2.650 | 1.067 | 1.929 | ||
| HiRAE | 0.023 | 7.905 | 14.722 | – | – | ||
| Model | rFID | gFID | |
|---|---|---|---|
| Selected 7 layers | |||
| RAEv2 | 0.299 | 1.060 | 2.170 |
| HiRAE-7 | 0.217 | 1.038 | 1.913 |
| All 24 layers | |||
| RAEv2 + DRoRAE | 0.065 | 1.551 | 3.375 |
| HiRAE-24 | 0.209 | 1.038 | 1.856 |
| Groups | rFID | Guided gFID |
|---|---|---|
| 2 | 3.293 | 29.385 |
| 3 | 3.363 | 28.354 |
| 4 | 3.457 | 31.321 |
| Groups | rFID | Guided gFID |
|---|---|---|
| 2 | 3.293 | 29.385 |
| 3 | 3.363 | 28.354 |
| 4 | 3.457 | 31.321 |
| Depth group | Group removal | High low |
|---|---|---|
| Shallow (0–7) | 5.650 | 0.215 |
| Middle (8–15) | 54.111 | 0.068 |
| Deep (16–23) | 147.750 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Layers | Norm cap | Dropout |
|---|---|---|---|
| Shallow | 0–7 | 0.025 | 0.50 |
| Middle | 8–15 | 0.075 | 0.25 |
| Deep | 16–23 | 0.150 | 0.10 |
| Setting | Stage 1 | Stage 2 |
|---|---|---|
| Epochs | 16 | 80 |
| Global batch size | 128 | 1024 |
| Gradient accumulation | 1 | 2 |
| EMA decay | 0.9978 | 0.9995 |
| Peak learning rate | ||
| Final learning rate |
| Setting | Pretraining | SFT |
|---|---|---|
| Hardware | H800 80GB | H800 80GB |
| Precision | BF16 mixed precision | BF16 mixed precision |
| Microbatch per GPU | 32 | 32 |
| Gradient accumulation | 4 | 4 |
| Global batch size | 1024 | 1024 |
| Optimizer updates | 100,000 | 2,850 |
| System | Generator | Parameters (M) |
|---|---|---|
| DiT / SD-VAE | DiT-XL/2 | 675 |
| SiT / SD-VAE | SiT-XL/2 | 675 |
| REPA | SiT-XL/2 | 675 |
| VA-VAE | LightningDiT-XL | 675 |
| REPA-E / E2E-VAE | SiT-XL/2 + REPA | 675 |
| RPiAE | LightningDiT | 675 |
| System | Unguided gFID | Guided gFID | Guided setting |
| RAE | 2.160 | 1.740 | AG |
| DRoRAE | 2.680 | 1.650 | AG |
| REPA-E / E2E-VAE | 3.460 | 1.670 | CFG |
| RPiAE | 2.250 | 1.510 | AG ∗ |
| FAE, timestep shift | 2.080 | 1.700 | CFG |
| DecQ, 8 queries | 1.800 | 1.330 | AG |
| Method | Epochs | gFID | |
|---|---|---|---|
| SiT-XL/2 | 800 | 2.120 | 8.440 |
| DDT-XL | 800 | 1.260 | 5.700 |
| SiT-XL/2 + REPA | 800 | 1.420 | 5.450 |
| LightningDiT | 800 | 1.420 | 4.570 |
| REG | 800 | 1.540 | 4.640 |
| REPA-E | 800 | 1.120 | 3.040 |
| Reference | PSNR | 95% interval | LPIPS | 95% interval |
|---|---|---|---|---|
| RAEv2 | ||||
| HiRAE-7 |
| Tokenizer | Sobel MSE | Laplacian MSE |
|---|---|---|
| RAEv2 | 1.830 | 43.378 |
| HiRAE-7 | 1.167 | 36.124 |
| HiRAE-24 | 0.816 | 32.645 |
| HiRAE-24 | HiRAE-7 | |
| Encoder layers used | All 24 | Selected 7 |
| Expert parameters (M) | 202 | 59 |
| Reconstruction rFID | 0.209 | 0.217 |
| PSNR (dB) | 26.377 | 25.068 |
| LPIPS | 0.043 | 0.049 |
| Guided gFID | 1.038 | 1.038 |
| Metric | DRoRAE-style fusion | HiRAE-24 | Reduction |
|---|---|---|---|
| 20 generator epochs | |||
| gFID | 2.244 | 2.242 | 0.1% |
| 5.270 | 2.533 | 51.9% | |
| 80 generator epochs | |||
| gFID | 1.551 | 1.038 | 33.1% |
| 3.375 | 1.856 | 45.0% | |
| Groups | Reconstruction | Unguided | IG=1.78 | ||||
|---|---|---|---|---|---|---|---|
| rFID | PSNR | LPIPS | gFID | IS | gFID | IS | |
| 2 | 3.293 | 22.417 | 0.1725 | 44.795 | 42.976 | 29.385 | 62.103 |
| 3 | 3.363 | 22.153 | 0.1754 | 44.055 | 43.216 | 28.354 | 63.796 |
| 4 | 3.457 | 22.065 | 0.1789 | 45.172 | 41.618 | 31.321 | 59.984 |
| Groups | Layers per group | Norm caps | Dropout probabilities |
|---|---|---|---|
| 2 | |||
| 3 | |||
| 4 |
| Group | Norm cap | LPIPS increase | Paired interval |
|---|---|---|---|
| Shallow (0–7) | 0.025 | 5.650 | [5.563, 5.737] |
| Middle (8–15) | 0.075 | 54.111 | [53.600, 54.626] |
| Deep (16–23) | 0.150 | 147.750 | [146.789, 148.661] |
| Group | High removal energy control | High low, equal norm | Paired interval |
|---|---|---|---|
| Shallow | 0.770 | 0.215 | [0.182, 0.248] |
| Middle | 4.760 | 0.068 | [ , 0.144] |
| Deep | 7.328 | [ , ] |
| Coordinates | Representation | High-frequency fraction | Same-class 10NN |
|---|---|---|---|
| Shared anchor | Deep anchor | 0.123 | 0.827 |
| Shared anchor | Fused representation | 0.208 | 0.815 |
| Raw | Deep anchor | 0.107 | 0.796 |
| Raw | Fused representation | 0.166 | 0.795 |
| Tokenizer | Gaussian | Low frequency | High frequency |
|---|---|---|---|
| RAEv2 | 8.385 | 7.191 | 9.414 |
| HiRAE-7 | 1.576 | 1.344 | 1.810 |
| HiRAE-24 | 1.791 | 1.547 | 2.043 |
| Coefficient log-SNR | RAEv2 | HiRAE-7 | HiRAE-24 |
|---|---|---|---|
| 0.459 | 0.617 | 0.473 | |
| 0.296 | 0.399 | 0.292 | |
| 0.171 | 0.224 | 0.165 | |
| 0 | 0.103 | 0.112 | 0.088 |
| 2 | 0.052 | 0.047 | 0.042 |
| 4 | 0.018 | 0.017 | 0.016 |