Variational autoencoders (VAEs) are a key part of modern text-to-image models, which generate images within their latent space. VAEs are known to disentangle the main factors of variation in the data, and color is known to be one of the most structured of these in natural images: decorrelating it yields one luminance axis and two opponent-color axes. Color should therefore be expected to emerge as a distinct factor in the VAE latent space. Yet how these latent spaces represent color remains largely unexplored. In this work, we show that the VAEs of text-to-image models share a color subspace aligned with brightness and opponent-colors. Through a linear approximation of the encoder and targeted latent steering, we find this subspace consistently across a broad range of VAEs, from SD1.5 to FLUX.2 and Z-Image. Building on this characterization, we propose three applications: ColorTuning, which achieves state-of-the-art in precise numerical color generation on the fine-grained CSS3/X11 system of GenColorBench, saturation control, to adjust the global chromatic intensity, and color transfer, to change the palette to match a reference. The code and models are publicly available at https://julian075.github.io/Color_Subspace/
Figures & tables
Figure 1: A color space hides in plain sight in the latent space of an autoencoder. Train a VAE with three latent channels on natural images, encode an image (left), add a constant offset δ to one channel, and decode: the image changes color while preserving its content. Sweeping two channels (right; ch2 across columns, ch3 across rows) spans an opponent-color plane and all intermediate mixtures, while the remaining channel (bottom row) controls brightness.
Figure 2: Image modifications through traverse channel variations across generative models: (a) Flux, (b) Flux 2, (c) Stable Diffusion 1.5, (d) SDXL, and (e) Stable Diffusion 3.
Figure 3: Texture-color variations through traverse latent channel modifications in different VAEs.
Figure 4: The first singular vectors of the linear approximation of VAEs and their corresponding CIE chromaticity directions across generative models: Flux.dev 2, Flux.dev 1, (c) SD 1.5, (d) SDXL, and SD 3. Each subplot visualizes the singular vectors ordered by their chromatic variance alongside their color content in the CIE chromatic diagram. We used 10616×16 patches from Tiny ImageNet . Appendix A.4 shows that this kind of vectors and opponent colors also emerge in other databases.
Table 1: Numerical color accuracy (%) on GenColorBench, on (a) the mini subset and (b) the full benchmark. The number of latent channels is shown in parentheses. Bold : best result for each backbone.
Figure 5: Qualitative results of ColorTuning . The generated examples demonstrate precise color control across three scenarios: applying distinct hues (top), generating subtle monochromatic variations within the same color family (middle), and pastel tones (bottom).
Figure 6: Further color control with the same color basis. Top: saturation reduction. From left to right, we intervene during generation to reduce the saturation of the final image. Bottom: color transfer, three examples. In each pair, the left image is the color reference and the right image is the generated result, whose colors are steered to match it.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Visual impact of individual latent channel shifts ( Δz ) across different AE.
Figure 8: Visual impact of individual latent channel shifts ( Δz ) across different KL-VAE.
Figure 9: Visual impact of individual latent channel shifts ( Δz ) across different VQ-VAE.
Figure 10: Latent channel shifts applied to KL-regularized VAEs with varying downsampling factors (f4, f8, f16).
Figure 11: Latent channel shifts across Vector-Quantized (VQ) VAEs.
Figure 12: Latent channel shifts in state-of-the-art foundation models. The top row compares the 4-channel autoencoders of SD1.5 and SDXL. The bottom row illustrates the expanded 16-channel latent spaces of SD3 and FLUX.1.
Figure 13: Extended latent channel shifts for the 32-channel FLUX2 autoencoder.
Figure 14: The first singular vectors of the linear approximation of VAEs and their corresponding CIE chromaticity directions across generative models used with 64×64 images: Flux.dev 2, Flux.dev 1, (c) SD 1.5, (d) SDXL, and SD 3. Each subplot visualizes the singular vectors ordered by their chromatic variance alongside their color content in the CIE chromatic diagram.
64×64 images
FLUX 2
FLUX
SD3
SDXL
SD1.5
AVERAGE
Lin. Enc. + Lin. Dec
Fract. Lin. Resp.
0.67
0.73
0.72
0.71
0.74
0.71 ± 0.02
Error No Linear
0.02
0.03
0.05
0.08
0.08
0.05 ± 0.02
Error Linear
0.11
0.09
0.09
0.09
0.09
0.09 ± 0.01
Lin. Enc. + NonLin. Dec
Appendix
Table 2: Linear approximation metrics across all models, configurations and datasets.
Figure 15: T2I models fail at numerical color precision.
Figure 16: Overview of our method. Top: characterizing the color basis of a VAE (Sec. 4.1 ). Each latent channel is perturbed inside the region of a set of reference objects, and the resulting shifts in CIELAB form a sensitivity matrix, whose SVD gives the luminance and opponent-color directions. This is computed once per VAE. Bottom: ColorTuning at inference (Sec. 4.2 ). The numerical code in the prompt is replaced by the nearest color name. At the gate step, the predicted clean latent z^0 is decoded, the object is segmented with SAM 3, and its mean color is measured. The MLP maps this color and the target to a displacement along the color basis, which is added inside the object mask over the remaining steps.
Model
Steps
CFG
Gate τgate
Temporal Schedule
FLUX.1-dev
28
3.5
0.50
Ascending
FLUX.2-dev
28
3.5
0.65
Ascending
SD3-Medium
28
7.0
0.75
Triangular
SD3.5-Medium
28
4.5
0.60
Ascending
SDXL-Base
30
5.0
0.40
Constant
Z-Image
28
4.0
0.60
Constant
Appendix
Table 3: Pipeline Specifications and Hyperparameters .
Figure 17: Qualitative results for saturation control.
Figure 18: Color transfer results. The first two columns display the target color palettes and their corresponding generated images. The final two columns present the color reference images alongside the resulting generated outputs.
Figure 19: Thurstone case V results of our user’s study. Values are z-scores. Error bars represent 95% confidence intervals —see Montag (2006) . Our method is statistically better than existing methods: ColorWave and NumColor.
Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffusion prior, but latent SDS often produces structured color artifacts and high-frequency texture noise. We identify a failure mode of latent SDS caused by VAE-induced pixel drift: the optimized image can move along pixel-space directions that are weakly constrained by the VAE encoder, so its latent representation remains clean and semantically meaningful while the image itself accumulates visible artifacts. We support this diagnosis with controlled 2D SDS experiments, VAE-only optimization, and a simplified analysis showing that encoder-like latent objectives can amplify image-space noise when the inverse mapping to pixels is underconstrained. Motivated by this observation, we propose PixSDS, a lightweight VAE-consistent gradient repair method. PixSDS decodes a latent SDS lookahead step and uses the decoded image as a clean direction for pixel-space optimization, reducing motion in VAE-inconsistent directions without retraining the diffusion model, changing the renderer, or replacing the SDS objective. Experiments in 2D optimization and text-to-3D generation show that PixSDS substantially reduces structured artifacts while preserving semantic content. Code is publicly available at https://sevashasla.github.io/pixsds-webpage/.
Most visual generative models compress images into a latent space before applying diffusion or autoregressive modelling. Yet, existing approaches such as VAEs and foundation model aligned encoders implicitly constrain the latent space without explicitly shaping its distribution, making it unclear which types of distributions are optimal for modeling. We introduce \textbf{Distribution-Matching VAE} (\textbf{DMVAE}), which explicitly aligns the encoder's latent distribution with an arbitrary reference distribution via a distribution matching constraint. This generalizes beyond the Gaussian prior of conventional VAEs, enabling alignment with distributions derived from self-supervised features, diffusion noise, or other prior distributions. With DMVAE, we can systematically investigate which latent distributions are more conducive to modeling, and we find that SSL-derived distributions provide an excellent balance between reconstruction fidelity and modeling efficiency, reaching gFID equals 3.2 on ImageNet with only 64 training epochs. Our results suggest that choosing a suitable latent distribution structure (achieved via distribution-level alignment), rather than relying on fixed priors, is key to bridging the gap between easy-to-model latents and high-fidelity image synthesis. Code is avaliable at https://github.com/sen-ye/dmvae.
Sen Ye, Jianning Pei, Mengde Xu +4
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · University of Chinese Academy of Sciences · Tencent +1
Text-to-image diffusion models generate images by gradually converting white Gaussian noise into a natural image. White Gaussian noise is well suited for producing diverse outputs from a single text prompt due to its absence of structure. However, this very property limits control over, and predictability of, specific visual attributes, as the noise is not human-interpretable. In this work, we investigate the characteristics of the input noise in diffusion models. We show that, although all frequencies in white Gaussian noise have comparable statistical energy, low-frequency components primarily determine the images global structure and color composition, while high-frequency components control finer details. Building on this observation, we demonstrate that simple manipulations of the low-frequency noise using low-frequency image priors can effectively condition the generation process to reconstruct these low-frequency visual cues. This allows us to define a simple, training-free method with minimal overhead that steers overall image structure and color, while letting high-frequency components freely emerge as fine details, enabling variability across generated outputs.