Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive. Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens. Training the pooling operator jointly with the RGB decoder preserves the standard two-stage RAE procedure without a separate feature auto-encoder. On ImageNet-256, 4x token compression retains comparable generation quality under internal guidance, while 16x compression trades some quality for greater efficiency. At a fixed budget of 100 sampling steps, latent-sampling throughput increases by 3.7x and 9.0x, respectively, relative to the unpooled baseline. Classification and dense prediction evaluations show that comparable guided generation quality can coexist with weaker performance on other tasks.
Figures & tables
Figure 1: Generation quality versus latent-sampling throughput on an NVIDIA H100 NVL at batch size 128. PoolDINO and the RAEv2 reference use Internal Guidance (IG) with 100 Euler steps. Red points use 80 training epochs; gold points use extended training (180 epochs for 2×2 and 300 for 4×4 ). Diamonds pair published FIDs with measured sampling rates. Appendix B details the measurement protocols and shows that reducing to 50 sampling steps doubles throughput while maintaining comparable generation quality.
Spatial pooling
Tokens
rFID ↓
sFID ↓
PSNR ↑
LPIPS ↓
RAEv2 baseline
256
0.32
2.29
22.74
0.15
Learned 2×2
64
0.36
2.61
22.25
0.17
Learned 2×4
32
0.39
2.76
21.88
0.18
Learned 4×2
32
0.39
2.82
21.85
0.18
Learned 4×4
16
0.41
2.99
21.45
0.19
Average 2×2
64
0.65
4.97
19.58
0.25
Table 1: Reconstruction on the ImageNet-1K 256×256 validation set, grouped by pooling method. All RGB decoders process 256 tokens after repetition.
Figure 3: Curated samples across spatial compression levels, with all generators trained for 80 epochs. Samples are selected independently across columns. Sampling uses 100 Euler steps and each model’s selected IG-only scale. Appendix G provides randomly selected samples without quality filtering.
Unguided
IG
IG + CFG
Spatial pooling
Tokens
FID ↓
IS ↑
FID ↓
IS ↑
FID ↓
IS ↑
RAEv2 baseline
256
1.53
226.63
1.08
262.00
1.07
267.15
Learned 2×2
64
3.00
184.49
1.09
250.36
1.07
268.50
Learned 2×4
32
4.91
161.42
1.19
249.58
1.16
273.06
Learned 4×2
32
4.97
159.67
1.21
249.96
1.19
273.98
Learned 4×4
16
8.16
132.67
1.44
247.07
1.35
282.29
Table 2: ImageNet-256 generation at 100 Euler steps. Models train for 80 epochs. We report generation FID and Inception Score (IS; Salimans et al. 2016 ). Guided columns report the lowest-FID settings from our guidance scale sweeps; the corresponding scales are in Table 18 . Best and second-best metrics in each column are bold and underlined.
Unguided
IG
Pooling
Epochs
FID ↓
IS ↑
FID ↓
IS ↑
1×1 (unpooled)
80
1.53
226.63
1.08
262.00
2×2
80
3.00
184.49
1.09
250.36
2×2
180
2.78
192.37
1.05
272.00
4×4
80
8.16
132.67
1.44
247.07
4×4
300
7.19
146.13
1.29
258.04
Table 3: Effect of extended training, evaluated with 100 Euler steps at the guidance settings in Table 18 . The unpooled 80-epoch model is included as a reference; bold marks the stronger result within each pooling window. Appendix A.2 gives longer runs and finer IG sweeps.
Figure 4: Curated comparisons of 80-epoch and extended training. Class conditions and initial noise are matched within each pair. All samples use 100 Euler steps and IG alone: scales 1.75/2.00 for the 80/180-epoch 2×2 models, and 2.75 for both 4×4 models.
ImageNet-1K
ADE20K
NYUv2
Spatial pooling
Window
Linear top-1 (%) ↑
mIoU (%) ↑
AbsRel ↓
Unpooled
1×1
85.31
46.86
0.0817
Learned
2×2
83.47
39.38
0.0895
Average
2×2
85.33
42.59
0.0869
Learned
4×4
80.80
38.61
0.1022
Average
4×4
85.33
36.76
0.0997
Table 4: Classification and dense prediction from frozen representations. Dense task models train from scratch. Bold marks the stronger pooling rule at each compression rate.
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
Pooling window
Latent grid
Tokens
Compression
κ
1×1 (control)
16×16
256
1×
8
2×2
8×8
64
4×
4
2×4
8×4
32
8×
8
4×2
4×8
32
8×
8
4×4
4×4
16
16×
2
Appendix
Table 5: Latent grids. Compression is relative to the 256-token control.
Hyperparameter
Image decoder
Flow encoder
Flow decoder
Transformer blocks
28
28
2
Hidden size
1,152
1,440
2,048
MLP width
4,096
3,840
5,461
Attention heads
16
20
16
Time tokens
–
4
–
Class tokens
–
8
–
Appendix
Table 6: Model architecture hyperparameters.
Hyperparameter
Decoder Stage
Flow-Model Stage
Training epochs
16
80
Training steps
40,032
100,080
Global batch size
512
1,024
Optimizer
AdamW
Muon (AdamW fallback)
Base learning rate
2×10−4
2×10−4
Learning rate schedule
Cosine decay to 2×10−5
Constant through epoch 25; linear decay to 2×10−5 at epoch 50
Appendix
Table 7: Optimization hyperparameters for both training stages.
Window
Tokens
Generator
Internal head
Repr. path
Decoder
Unguided generation
1×1
256
473.182G
2.887G
0.756G
224.216G
47.542T
2×2
64
128.653G
0.722G
1.292G
224.216G
13.090T
2×4
32
72.531G
0.361G
1.292G
224.216G
7.477T
4×4
16
44.612G
0.180G
1.292G
224.216G
4.685T
Appendix
Table 8: Estimated forward FLOPs per image, counting a multiply-add as two operations. Component columns report one evaluation. “Internal head” is the additional cost of returning the early clean-latent prediction. “Representation path” includes the dense-representation projection and pooling operation. Unguided generation counts 100 base generator evaluations and one image-decoder evaluation, excluding auxiliary heads and additional CFG evaluations.
180 epochs
222 epochs
IG scale
FID ↓
IS ↑
FID ↓
IS ↑
1.000
2.7805
192.37
2.7682
193.79
1.250
1.7069
218.79
1.7269
220.06
1.500
1.2237
240.43
1.2330
242.09
1.750
1.0603
257.87
1.0645
259.57
1.875
1.0462
265.43
1.0476
266.27
Appendix
Table 9: IG sweeps for 2×2 learned pooling at 180 and 222 epochs. Bold marks the lowest observed FID at each duration.
300 epochs
390 epochs
IG scale
FID ↓
IS ↑
FID ↓
IS ↑
1.000
7.1942
146.13
6.7782
150.30
1.250
4.7407
171.29
4.4667
176.07
1.500
3.1969
194.26
3.0450
198.28
1.750
2.2363
214.17
2.1802
218.03
2.000
1.7026
229.57
1.6962
233.37
Appendix
Table 10: IG sweeps for 4×4 learned pooling at 300 and 390 epochs. Bold marks the lowest observed FID at each duration.
Figure 6: Quality–throughput comparison with both guidance settings on an H100 NVL at batch size 128. Our 80-epoch models use IG alone (solid, filled) or CFG+IG (dashed, hollow), with FID evaluated at 100 Euler steps. Point labels give token compression; black circles denote the unpooled RAEv2 reference. Literature baselines and timing conventions match Figure 1 ; full-sampling throughput excludes image decoding.
Figure 7: Inception Score versus latent-sampling throughput on an H100 NVL at batch size 128. PoolDINO and the unpooled RAEv2 reference use IG alone at 100 Euler steps, including the extended-training checkpoints. IS is taken at the FID-selected guidance settings, not maximized separately.
FID ↓
Latent samples/s ↑
Pooling
Epochs
sIG
100 steps
50 steps
100 steps
50 steps
2×2
80
1.75
1.0890
1.1164
24.16±0.05
48.44±0.17
2×2
180
2.00
1.0525
1.0629
24.16±0.05
48.44±0.17
2×4
80
2.00
1.1913
1.2283
40.10±0.07
80.42±0.33
4×4
80
2.75
1.4442
1.4429
58.46±0.20
117.13±0.38
4×4
300
2.75
1.2882
1.2981
58.46±0.20
117.13±0.38
Appendix
Table 11: IG-only generation with 100 versus 50 Euler steps. Throughput is latent samples/s on an H100 NVL at batch size 128; ± reports timing standard deviation.
Figure 8: Generation quality versus latent-sampling throughput with 50-step sampling for PoolDINO, using IG alone on an H100 NVL at batch size 128. Red points show the 80-epoch models; gold points show extended training. The RAEv2 reference uses IG with 100 steps. Literature points, including FlatDINO, retain their published FIDs and measured throughput under their respective sampling protocols.
Figure 9: Sampling-step ablation for the 4×4 generator trained for 300 epochs. Gold shows FID (left axis); red shows IS (right axis). The IG scale is fixed at 2.75 rather than retuned for each step count.
Figure 10: Denoiser forward-pass speed on a single GPU, using the flow-model architecture in Table 6 . Legend entries give pooling windows and token compression. Each point is the mean of 30 measurements averaging 20 forward calls; error bars show one standard deviation and are mostly smaller than the markers. Benchmarks use BF16 computation and include host dispatch and synchronization, but not end-to-end image generation.
Figure 11: Generator forward–backward throughput on a single GPU. Points show mean iterations/s across 30 measurements of 20 iterations. Timing includes the forward pass, main-prediction MSE loss, backward pass, and dispatch/synchronization; it excludes optimizer updates, auxiliary losses, and the encoder/decoder. The uncompressed model runs out of memory at batch size 256 on both GPUs.
k-NN
Linear probe
Window
Learned
Average
PCA
Random
Learned
Average
PCA
Random
Uncompressed z
76.03
85.31
2×2
72.16
76.00
76.67
75.78±0.06
83.47
85.33
84.35
84.63±0.06
2×4
68.50
76.00
76.58
75.77±0.03
82.35
85.33
84.10
84.50±0.04
4×2
68.17
76.00
76.59
75.68±0.05
82.35
85.34
84.12
84.43±0.07
4×4
63.22
76.00
76.39
75.68±0.04
80.80
85.33
83.77
84.21±0.04
Appendix
Table 12: ImageNet-1K k-NN and linear-probe top-1 accuracy (%). The uncompressed RAEv2 representation z is included as a reference. Best and second-best methods at each compression rate are bold and underlined.
Spatial pooling
Best epoch
Val. loss ↓
mIoU ↑
Mean acc. ↑
Pixel acc. ↑
Uncompressed
75
0.671
46.860
58.060
81.260
Learned 2×2
75
0.807
39.380
50.100
77.480
Average 2×2
65
0.735
42.591
53.477
78.905
Learned 4×4
80
0.863
38.610
48.850
77.780
Average 4×4
75
0.866
36.758
46.969
75.715
Appendix
Table 13: ADE20K semantic segmentation with task models trained from scratch. Accuracy metrics are percentages.
Spatial pooling
Best epoch
AbsRel ↓
RMSE ↓
Log RMSE ↓
SiLog ↓
δ1↑
δ2↑
δ3↑
Uncompressed
12
0.082
0.409
0.122
0.122
94.170
99.030
99.780
Learned 2×2
21
0.090
0.455
0.134
0.134
92.780
98.700
99.720
Average 2×2
14
0.087
0.434
0.131
0.131
93.200
98.720
99.730
Learned 4×4
27
0.102
0.490
0.148
0.148
90.440
98.260
99.610
Average 4×4
15
0.100
0.486
0.149
0.149
90.580
98.060
99.560
Appendix
Table 14: NYUv2 monocular depth estimation with task models trained from scratch. The δ metrics are percentages.
Spatial pooling
Tokens
Compression
Best epoch
Val. loss ↓
mIoU ↑
Mean acc. ↑
Pixel acc. ↑
Uncompressed
256
1×
80
0.694
47.810
58.850
81.770
Learned 2×2
64
4×
60
0.709
45.380
56.420
81.010
Learned 4×4
16
16×
80
0.788
42.500
53.240
79.960
Appendix
Table 15: ADE20K semantic segmentation with task models initialized from the corresponding RGB decoder. Accuracy metrics are percentages.
Spatial pooling
Best epoch
AbsRel ↓
RMSE ↓
Log RMSE ↓
SiLog ↓
δ1↑
δ2↑
δ3↑
Uncompressed
14
0.081
0.411
0.123
0.123
94.330
98.980
99.750
Learned 2×2
27
0.089
0.456
0.133
0.133
92.800
98.720
99.700
Learned 4×4
24
0.101
0.495
0.147
0.147
90.610
98.220
99.610
Appendix
Table 16: NYUv2 monocular depth estimation with task models initialized from the corresponding RGB decoder. The δ metrics are percentages.
Unguided
CFG
IG
CFG + IG
Variant
FID ↓
IS ↑
FID ↓
IS ↑
FID ↓
IS ↑
FID ↓
IS ↑
1×1
1.53
226.63
1.45
237.80
1.08
262.00
1.07
267.15
2×2
3.00
184.49
1.99
228.10
1.09
250.36
1.07
268.50
2×4
4.91
161.42
2.33
238.80
1.19
249.58
1.16
273.06
4×2
4.97
159.67
2.41
237.63
1.21
249.96
1.19
273.98
4×4
8.16
132.67
2.87
249.72
1.44
247.07
1.35
282.29
Appendix
Table 17: ImageNet-256 generation at 100 Euler steps. Models train for 80 epochs, except the extended-training 2×2 (180 epochs) and 4×4 (300 epochs) variants. Guided columns use the settings in Table 18 ; Appendix A.2 separately reports longer runs and finer IG sweeps. – denotes an unevaluated setting. Best and second-best metrics in each column are bold and underlined.
CFG
IG
CFG + IG
Variant
sCFG
sIG
sCFG
sIG
1×1
3.00
1.75
2.50
1.78
2×2
3.00
1.75
1.75
1.75
2×4
3.00
2.00
1.50
2.00
4×2
3.00
2.00
1.50
2.00
4×4
3.00
2.75
1.75
2.25
Appendix
Table 18: Guidance scales selected for Tables 2 , 3 , and 17 . For CFG alone, sIG=1 ; for IG alone, sCFG=1 . Both scales are one for unguided sampling.
Method
Euler steps
Guidance scale
Guidance interval
Neutral value
Unguided
100
–
–
CFG 1
Internal guidance
100
Swept
[0,0.9]
1
Encoder-reconstruction guidance
100
Swept
[0,1]
0
Appendix
Table 19: Sampling settings for the guidance ablations. Intervals use noise-to-data time: t=0 is noise and t=1 is clean data.
Figure 12: Isolated guidance-scale sweeps on ImageNet-256. Top: FID on a logarithmic scale. Bottom: Inception Score. Missing points in the corrected 2×2 sweeps are left blank.
Encoder-reconstruction guidance
Internal guidance
Variant
FID ↓
IS ↑
FID ↓
IS ↑
2×2
1.06
272.83
1.07
268.50
2×4
1.17
279.68
1.16
273.06
4×2
1.20
280.14
1.19
273.98
4×4
1.35
295.39
1.35
282.29
Appendix
Table 20: Metrics at the lowest-FID point in each joint guidance sweep for the 80-epoch models. Both auxiliary methods are combined with CFG.
Figure 13: Joint CFG and encoder-reconstruction guidance search. Squares show evaluated 100-step configurations; the black outline marks the lowest FID in each panel.
Figure 14: Inception Score for the joint CFG and encoder-reconstruction guidance search. The black outline marks the highest score in each panel.
Figure 15: Joint CFG and internal-guidance search. Squares show evaluated 100-step configurations; the black outline marks the lowest FID in each panel.
Figure 16: Inception Score for the joint CFG and internal-guidance search. The black outline marks the highest score in each panel.
Window
Average alignment
PCA alignment
Energy captured (% of PCA)
2×2
0.5000
0.2798
67.60
2×4
0.3530
0.2114
52.16
4×2
0.3501
0.2082
52.05
4×4
0.2268
0.1524
38.16
Appendix
Table 21: Learned pooling row-space analysis. Alignment scores are adjusted so that zero indicates the overlap expected between random subspaces of the same dimension and one indicates identical subspaces. Energy captured is the squared validation feature energy preserved by the learned pooling operator’s row space (after centering with the training mean), expressed as a percentage of that captured by PCA with the same output dimension.
Average overlap
PCA overlap
Window
Random
Raw
Adjusted
Raw
Adjusted
Energy (% PCA)
2×2
0.2500
0.6250
0.5000
0.4598
0.2798
67.60
2×4
0.1250
0.4339
0.3530
0.3100
0.2114
52.16
4×2
0.1250
0.4313
0.3501
0.3072
0.2082
52.05
4×4
0.0625
0.2751
0.2268
0.2054
0.1524
38.16
Appendix
Table 22: Complete learned pooling row-space results. “Random” is the expected raw overlap of random subspaces of the same dimension. Energy captured is reported as a percentage of rate-matched PCA.
Figure 17: Unpooled RAEv2 reference ( 1×1 ), trained for 80 epochs, with sIG=1.75 and no CFG. Samples are randomly selected without quality filtering.
Figure 18: PoolDINO with 2×2 learned pooling ( 4× token compression), trained for 80 epochs, with sIG=1.75 and no CFG. Samples are randomly selected without quality filtering.
Figure 19: PoolDINO with 2×4 learned pooling ( 8× token compression), trained for 80 epochs, with sIG=2.00 and no CFG. Samples are randomly selected without quality filtering.
Figure 20: PoolDINO with 4×4 learned pooling ( 16× token compression), trained for 80 epochs, with sIG=2.75 and no CFG. Samples are randomly selected without quality filtering.
Figure 21: PoolDINO with 2×2 learned pooling and extended training ( 4× token compression), trained for 180 epochs, with sIG=2.00 and no CFG. Samples are randomly selected without quality filtering.
Figure 22: PoolDINO with 4×4 learned pooling and extended training ( 16× token compression), trained for 300 epochs, with sIG=2.75 and no CFG. Samples are randomly selected without quality filtering.
Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insights which simplify and improve RAE. First, we study a generalized formulation where the representation is defined as sum of the last k encoder layers rather than solely the final layer. This simple change greatly improves reconstruction without encoder finetuning or specialized data (e.g., text, faces). Second, we study the prevalent assumption that RAE (using pretrained representation as encoder) replaces representation alignment (REPA), which distills the same representation to intermediate layers instead. Through large-scale empirical analysis, we uncover a surprising finding: RAE and REPA exhibit complementary working mechanisms, allowing the same representation to be used as both encoder and target for intermediate diffusion layers. Finally, the original RAE struggles with classifier-free guidance (CFG) and requires training a second, weaker diffusion model for AutoGuidance (AG). We show that REPA itself can be viewed as x-prediction in RAE latent space. By simply re-parameterizing the output of the DiT model, it can provide guidance for "free". Overall, RAEv2 leads to more than 10x faster convergence over the original RAE, achieving a state-of-the-art gFID of 1.06 in just 80 epochs on ImageNet-256. On FDr6, RAEv2 achieves a state-of-the-art 2.17 at just 80 epochs compared to the previous best 3.26 (800 epochs) without any post-training. This motivates EPFID@k (epochs to reach unguided gFID < k) as a measure of training efficiency. RAEv2 attains an EPFID@2 of 35 epochs, versus 177 for the original RAE. We also validate our approach across diverse settings for text-to-image generation and navigation world models, showing consistent improvements. The code is available at https://raev2.github.io.
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off: shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling preserves the full-layer latent mean in expectation while explicitly penalizing sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. Applying FuseReg to both stages also reduces unguided gFID by 29% on DiT-Base. The reconstruction and generation benefits also extend to other encoder families. FuseReg narrows the reconstruction-generation gap without additional training cost or architectural changes.
Hongyang Du, Yunfei Xie, Junjie Ye +13
USC PSI Lab · Brown University · Rice University +4
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.
Xuanyu Zhu, Yan Bai, Yang Shi +5
Peking University · Agibot Research · Tsinghua University +1