FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
Authors: Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, +8 more
Organizations: USC PSI Lab · Brown University · Rice University · University of Aberdeen · University of Notre Dame · University of Maryland, College Park · University of Pennsylvania
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off: shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling preserves the full-layer latent mean in expectation while explicitly penalizing sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. Applying FuseReg to both stages also reduces unguided gFID by 29% on DiT-Base. The reconstruction and generation benefits also extend to other encoder families. FuseReg narrows the reconstruction-generation gap without additional training cost or architectural changes.
Figures & tables
Figure 2: Cosine-similarity maps of decoder intermediate-block features under different layer fusions. For three query patches (stars), we visualize patch-wise cosine similarity at a decoder intermediate block, for FuseReg (top) and RAEv2 (bottom), fed three fusions: K=23 , K=7 , and the single layer ℓ11 ; the rightmost column is the reconstruction. FuseReg keeps clean, stable spatial structure across all fusions, whereas RAEv2 degrades away from its K=23 training fusion.
fusion k=7
fusion k=23
fusion ℓ11
Decoder
PSNR ↑
SSIM ↑
rFID ↓
PSNR ↑
SSIM ↑
rFID ↓
PSNR ↑
SSIM ↑
rFID ↓
RAEv2 K=7
22.58
0.626
0.30
18.37
0.531
1.62
17.50
0.491
6.35
RAEv2 K=23
12.51
0.369
16.10
27.04
0.806
0.18
14.31
0.447
11.21
FuseReg p=.95
23.77
0.678
0.60
27.52
0.826
0.42
25.13
0.735
0.45
Table 1: One FuseReg decoder matches or outperforms any deterministic fusion. Light-blue cells mark the fusion(s) seen during training: each RAEv2 decoder is trained on one deterministic aggregate ( k=7 or k=23 ). The FuseReg decoder is trained across all nontrivial fusions.
Figure 3: Fixed generator, swapped decoders. A single reproduced RAEv2 DiT ( pdit=0 ) produces one fixed set of latents in the k=23 or k=7 fusion space; each point renders those same latents with a decoder trained at FuseReg rate pdec (horizontal axis), so changes in the rendered images arise solely from decoder replacement. Orange : gFID ↓ (left axis); blue : IS ↑ (right axis). The pdec=0 point is the plain RAEv2 decoder. gFID/IS over 50k samples; guidance is the official RAEv2 internal guidance (scale 1.78 ) ( Singh et al., 2026 ) ; k=7 panels use a log gFID scale. Full numbers are in Table D.1 .
Table 2: FuseReg across two stages and generator scales. Unguided generation on class-conditional ImageNet-256, 40 epochs, 50k samples. Rows: generator rate pdit ; columns: decoder rate pdec . The green box at (0,0) is the reproduced RAEv2 fixed-fusion baseline. Shading diverges from that baseline ( Blue = better, Orange = worse). Bold : best per panel.
Figure 4: FuseReg distributes reconstruction across layers. Frozen DINOv3, K=23 : the RAEv2 decoder and FuseReg ( pdec=0.95 ), evaluated on the same 10,000 held-out ImageNet-256 images with the fixed ℓ23 token-mean surrogate. Left: removing individual layers yields a flatter PSNR-drop profile for FuseReg. Right: individual layers yield stronger reconstructions; dashed lines mark full-fusion PSNR.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Epochs
Drop rates
Decoder
16
pdec∈{0,0.05,0.1,0.3,0.5,0.7,0.9,0.95}
DiT-Base
40
pdit,pdec∈{0,0.05,0.1,0.3,0.5,0.7,0.9}
DiT-XL
40
pdit∈{0,0.5,0.7,0.9} ; decoder rates as above
Appendix
Table B.1: Training budgets and FuseReg rate grids. Decoder and generator rates are varied independently. In the generator studies, each trained DiT is evaluated with the corresponding decoder-rate grid.
Fusion k=7
Fusion k=23
Single layer ℓ11
pdec
PSNR ↑
SSIM ↑
rFID ↓
PSNR ↑
SSIM ↑
rFID ↓
PSNR ↑
SSIM ↑
rFID ↓
0 (RAEv2)
12.54
0.372
16.108
27.10
0.808
0.189
15.81
0.503
3.392
0.05
19.78
0.518
2.826
28.59
0.853
0.273
20.76
0.557
1.881
0.1
20.52
0.550
1.718
28.58
0.852
0.285
21.33
0.586
1.209
0.3
22.05
0.612
0.762
28.48
0.849
0.294
22.90
0.657
0.576
0.5
22.93
0.646
0.655
28.43
0.848
0.294
23.98
0.695
0.522
Appendix
Table C.1: FuseReg controls decoder robustness across layer fusions. Each row is one decoder trained for 16 epochs with rate pdec and evaluated on 50k ImageNet-256 images using three input fusions. All runs retain the GAN and classifier-surrogate terms. The pdec=0 row is our reproduced full-fusion baseline. Bold / underline : best/second-best within each metric and fusion.
Figure C.1: One FuseReg decoder supports both sparse and single-layer fusions. Each row compares one image under the k=7 layers (blue) and ℓ11 (orange). RAEv2 labels denote the fixed training fusion: K=7 is matched in the blue block; both baselines are shifted in the orange block. The same FuseReg decoder preserves image structure in both settings. Insets show PSNR (dB).
Figure C.2: Subset diagnostics characterize reconstruction quality, layer contributions, and diminishing gains. Frozen DINOv3 features, K=23 ; gray denotes RAEv2 and orange denotes FuseReg. (a) Reconstruction PSNR versus subset size; lines show means and bands the 10th–90th percentiles over evaluated image–subset pairs. (b) Monte-Carlo Shapley estimates attribute PSNR gains to individual layers across sampled permutations. (c,d) Differences between adjacent subset sizes, measured as increases in mean PSNR and reductions in mean MSE, respectively. FuseReg achieves strong reconstruction from small subsets and reaches diminishing returns earlier while broadly retaining the ordering of layer contributions.
No guidance
Internal guidance (scale 1.78 )
DiT k=23
DiT k=7
DiT k=23
DiT k=7
Decoder pdec
gFID ↓
IS ↑
gFID ↓
IS ↑
gFID ↓
IS ↑
gFID ↓
IS ↑
0.0 (RAEv2)
3.01
206.8
27.73
97.7
1.25
263.8
16.35
148.7
0.05
3.15
209.5
6.24
185.2
1.48
255.7
4.62
213.5
0.1
3.16
209.4
4.30
196.2
1.49
256.7
3.15
226.1
0.3
2.92
211.6
2.68
209.4
1.41
257.4
1.76
242.3
Appendix
Table D.1: FuseReg improves generation with a fixed generator. Within each fusion and guidance setting, every decoder renders the same 50k latents from a fixed RAEv2 DiT ( pdit=0 ). The gray row uses the reproduced fixed- k=23 decoder; other rows vary pdec over the same K=23 layer pool. Internal guidance uses scale 1.78 ( Singh et al., 2026 ) . Bold : best in each column.
Table D.2: FuseReg rate preferences under internal guidance on DiT-XL. Class-conditional ImageNet-256, K=23 , hidden size 1440, 40 training epochs, and 50k samples with guidance scale 1.78 . Rows vary pdit and columns vary pdec . Colors encode metric values on a shared scale within each panel: blue denotes better values (lower gFID or higher IS), and orange denotes worse values. Bold : best in each panel.
SigLIP2-L ( K=23 )
Fusion k=7
Fusion k=23
Single layer ℓ11
pdec
PSNR ↑
SSIM ↑
rFID ↓
PSNR ↑
SSIM ↑
rFID ↓
PSNR ↑
SSIM ↑
rFID ↓
0 (RAEv2)
13.05
0.288
107.689
27.92
0.829
0.295
11.27
0.193
198.638
0.3
21.73
0.594
0.809
27.87
0.827
0.325
22.38
0.623
0.744
0.6
22.78
0.635
0.947
27.80
0.825
0.452
23.58
0.667
0.860
0.9
23.30
0.651
0.999
27.03
0.803
0.522
24.17
0.684
0.815
Appendix
Table E.1: Reconstruction robustness of FuseReg across encoder families. Each row is one decoder trained for 16 epochs at rate pdec and evaluated on 50k ImageNet-256 images with sparse, full, and single-layer inputs. The encoder headings separate the two families and their different layer subsets; every fusion includes the same last-layer classifier surrogate for that encoder. Gray rows mark the full-fusion baselines ( pdec=0 ). Bold / underline : best/second-best.
Table E.2: FuseReg at both stages across encoder families. Unguided generation on class-conditional ImageNet-256 with frozen SigLIP2-L ( K=23 , top) and EUPE-B ( K=11 , bottom) encoders, 16-epoch decoders, and 40-epoch DiT-Base generators; 50k samples and 50 Euler steps. Rows vary pdit ; columns vary pdec , with sampled latents fixed within each row. Each green box marks that encoder’s fixed-fusion baseline. Shading is relative to this baseline ( Blue = better, Orange = worse), with intensity scaled within each panel. Bold : best per panel.
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.
Xuanyu Zhu, Yan Bai, Yang Shi +5
Peking University · Agibot Research · Tsinghua University +1
Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insights which simplify and improve RAE. First, we study a generalized formulation where the representation is defined as sum of the last k encoder layers rather than solely the final layer. This simple change greatly improves reconstruction without encoder finetuning or specialized data (e.g., text, faces). Second, we study the prevalent assumption that RAE (using pretrained representation as encoder) replaces representation alignment (REPA), which distills the same representation to intermediate layers instead. Through large-scale empirical analysis, we uncover a surprising finding: RAE and REPA exhibit complementary working mechanisms, allowing the same representation to be used as both encoder and target for intermediate diffusion layers. Finally, the original RAE struggles with classifier-free guidance (CFG) and requires training a second, weaker diffusion model for AutoGuidance (AG). We show that REPA itself can be viewed as x-prediction in RAE latent space. By simply re-parameterizing the output of the DiT model, it can provide guidance for "free". Overall, RAEv2 leads to more than 10x faster convergence over the original RAE, achieving a state-of-the-art gFID of 1.06 in just 80 epochs on ImageNet-256. On FDr6, RAEv2 achieves a state-of-the-art 2.17 at just 80 epochs compared to the previous best 3.26 (800 epochs) without any post-training. This motivates EPFID@k (epochs to reach unguided gFID < k) as a measure of training efficiency. RAEv2 attains an EPFID@2 of 35 epochs, versus 177 for the original RAE. We also validate our approach across diverse settings for text-to-image generation and navigation world models, showing consistent improvements. The code is available at https://raev2.github.io.
Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quality. However, existing methods universally extract features from only the last encoder layer, discarding the rich hierarchical information distributed across intermediate layers. We show that low-level visual details survive in the last layer merely as attenuated residuals after multiple layers of semantic abstraction, and that explicitly fusing multi-layer features can substantially recover this lost information. We propose DRoRAE (Depth-Routed Representation AutoEncoder), a lightweight fusion module that adaptively aggregates all encoder layers via energy-constrained routing and incremental correction, producing an enriched latent compatible with a frozen pretrained decoder. A three-phase decoupled training strategy first learns the fusion under the implicit distributional constraint of the frozen decoder, then fine-tunes the decoder to fully exploit the enriched representation. On ImageNet-256, DRoRAE reduces rFID from 0.57 to 0.29 and improves generation FID from 1.74 to 1.65 (with AutoGuidance), with gains also transferring to text-to-image synthesis. Furthermore, we uncover a log-linear scaling law (R2=0.86) between fusion capacity and reconstruction quality, identifying \textit{representation richness} as a new, predictably scalable dimension for visual tokenizers analogous to vocabulary size in NLP.
Xuanyu Zhu, Yan Bai, Yang Shi +4
Peking University · Meituan Inc · Tsinghua University +1