cs.CVOct 3, 2025

What Drives Compositional Generalization in Visual Generative Models? The Importance of Continuous Training Objectives

Authors: Karim Farid, Rajat Sahay, Yumna Ali Alnaggar, Simon Schrodi, Volker Fischer, Cordelia Schmid, Thomas Brox

Organizations: University of Freiburg · Bosch Center for Artificial Intelligence · Inria, École Normale Supérieure, CNRS, PSL Research University

Abstract

Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of which design choices critically determine compositional generalization in image and video generation. By isolating independent design axes, we identify two key factors strongly associated with compositional success: (i) whether the training objective operates on a discrete or continuous distribution, and (ii) the completeness of conditioning information about constituent factors during training. We also show that relaxing the discrete loss with an auxiliary continuous latent objective can partially recover compositional performance in discrete models like MaskGIT. Our findings, corroborated by diverse compositional tasks and preliminary evidence in world models and LLMs, motivate a shift toward continuous objectives for compositional generalization.

Figures & tables

Appendix figures & tables32 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 22, 2026cs.LG

Catastrophic Compositional Generation: Why Vanilla Diffusion Models Fail to Extrapolate

The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to produce samples from compositionally-defined target distributions such as a geometric combination of the source distributions. In this work, we argue that this task is often infeasible for vanilla conditional diffusion models: we conjecture that no inference-time technique can efficiently produce samples from the target distribution in certain well-motivated settings. This idea is supported by theory-guided generalization arguments and carefully-designed experiments on both synthetic and realistic data. In particular, while recent methods such as Feynman-Kac correction reduce inference-time approximation error, our results show that score estimation error has a more catastrophic effect on performance when the target distribution is out-of-distribution with respect to the sources, highlighting the need for a different approach to this task.
May 14, 2026cs.CV

Compositional Video Generation via Inference-Time Guidance

Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relations between entities, attributes, actions, and motion directions. We hypothesize that these failures need not be addressed by retraining the generator, but can instead be mitigated by steering the denoising process using the model's own internal grounding signals. We propose \textbf{CVG}, an inference-time guidance method for improving compositional faithfulness in frozen text-to-video models. Our key observation is that cross-attention maps already encode how prompt concepts are grounded across space and time. We train a lightweight compositional classifier on these attention features and use its gradients during early denoising steps to steer the latent trajectory toward the desired composition. Built on a frozen VLM backbone, the classifier transfers across semantically related composition labels rather than relying only on narrow category-specific features. CVG improves compositional generation without modifying the model architecture, fine-tuning the generator, or requiring layouts, boxes, or other user-supplied controls. Experiments on compositional text-to-video benchmarks show improved prompt faithfulness while preserving the visual quality of the underlying generator.
Sep 27, 2026cs.LG

OOD Generalization as a Bifurcation Problem

Systematic out-of-distribution (OOD) generation remains a critical bottleneck for continuous-time generative models. While standard joint classifier-free guidance (CFG) routinely fails to synthesize unobserved concept combinations, exact decomposed scoring generalizes robustly at the cost of severe computational overhead. In this work, we reveal that compositional binding is not a uniform process but a highly localized phase transition. We identify the semantic bifurcation window - the precise temporal interval where joint and decomposed vector fields meaningfully diverge. Exploiting this dynamic, we propose surgical guidance, a hybrid sampling strategy that restricts exact multi-pass scoring strictly to this critical window. On an OOD bi-digit MNIST testbed, surgical guidance achieves state-of-the-art compositional fidelity at a fraction of the inference cost, yielding a +5.3% absolute improvement in pairwise accuracy over the joint baseline by intervening during just the first 15% of the diffusion trajectory. Furthermore, our empirical analysis uncovers a fundamental topological divide: diffusion models (SDEs) force conceptual resolution immediately at peak noise, whereas Conditional Flow Matching (ODEs) delays structural binding until intermediate features emerge, establishing a new temporal framework for accelerating large-scale generative decoding.