cs.CVSep 23, 2026

On the Diffusibility of High-Dimensional Latents

Authors: Chao FengZhiyang XuBowei ChenYuanjun XiongXiyao WangJui-Hsien WangRichard ZhangZhe Lin+2 more

Abstract

Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction recovers such details. However, perhaps counterintuitively, this procedure reduces the effective dimensionality of the resulting representation, and the altered geometry has downstream effects on generation. Specifically, we show that using the standard velocity prediction in flow matching in this high-dimensional space requires the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient. This motivates using the clean data parameterization (x0\boldsymbol{x}_{0}-prediction) instead, which focuses learning on the underlying signal manifold. Across experiments with multiple strong-reconstruction encoders, we show that x0\boldsymbol{x}_{0}-prediction consistently improves text-to-image generation performance.

Explore similar work

CardsList
  1. Improved Baselines with Representation Autoencoders

    May 18, 2026Jaskirat Singh, Boyang Zheng, Zongze Wu +3Masked AutoencodersImagenet