cs.CVSep 28, 2026

Generative Residual Factorization

Authors: Letian Gong, Yuzhou Hong

Organizations: Zhejiang University of Science and Technology · Zhejiang Sci-Tech University

Abstract

Under a shared-factor model, the conditional law of the next image patch factors into a posterior over the shared scene factor and a residual kernel given that factor. A sufficient statistic of the past replaces the raw past in the posterior and does not replace the kernel. The conditional entropy splits into residual entropy, which no observation of the factor can remove, and a posterior term, which a better representation of the past can remove. Next-embedding prediction is a directional likelihood on a shallow map, so the fiber of that map is unidentified and a constant embedding remains a minimizer. The same split is an equality in a scalar Gaussian model, evaluated in closed form.

Figures & tables

Explore similar work

Sep 25, 2026cs.CV

Conditional Predictive Sufficient Statistics for Visual Representation Learning

A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.
Aug 5, 2026cs.LG

Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models

Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information. We show that this interpretation is incorrect. Matching only the marginal distribution places no constraint on the class-conditional distributions, allowing the latent style to remain highly predictive of the label despite appearing perfectly Gaussian in aggregate. We derive an exact decomposition showing that this mismatch is one of four conditions required for factorized sampling, and demonstrate that eliminating it is necessary but not sufficient to obtain the intended factorization. Empirically, our case-study model and four representative latent baselines achieve near-zero global MMD while still allowing a linear probe to recover class labels with 74%--100% accuracy (10% chance level). Our model reaches 99.15% clustering accuracy, whereas externally evaluated class-conditional generation succeeds only 16% of the time. This leakage remains under six independent perturbations involving model capacity, curriculum, prior geometry, and supervision across two datasets. Four mitigation strategies reduce probe accuracy to 21%--46%, although they leave within-class dependence largely unchanged. A post-hoc conditional prior improves externally evaluated class generation to 0.97 on MNIST without retraining but reaches only 0.41 on CIFAR-10, while an empirical style bank achieves 0.88 on CIFAR-10. These results demonstrate that no divergence computed solely on the marginal distribution of the style latent can certify independence from class labels, and that reporting marginal statistics alone does not verify the property commonly claimed in factorized generative models.
May 16, 2025cs.CV

Conditioning Residuals for Diffusion Models via Representation Feedback

Diffusion models now serve as a common foundation for multimedia generation, and useful intermediate representations emerge during their generative training. Standard architectures, however, propagate these representations through the main feature stream, without explicitly reintroducing their encoded semantics to later denoising layers. Meanwhile, such backbones already provide a conditioning pathway for global modulation by predefined inputs. This work examines whether this native pathway can also route internally inferred semantics as evolving, sample-dependent cues. We propose Conditioning Residuals, a lightweight feedback mechanism that converts aggregated features into residuals added to condition embeddings. By feeding back compact feature summaries, it provides adaptive generative guidance and encourages a tighter semantic bottleneck, without external encoders, auxiliary objectives, or sampling-time changes. It supports feedback at one or multiple depths in UNet and DiT backbones, with negligible overhead. Across diffusion formulations, backbone configurations, and datasets, experiments show consistent gains in generative performance, along with stronger representations in downstream linear probing and segmentation. Mechanistic analyses reveal improved generative training dynamics and reshaped feature structure, suggesting a grounded, generalizable way to enhance diffusion backbones from within.