Systematic out-of-distribution (OOD) generation remains a critical bottleneck for continuous-time generative models. While standard joint classifier-free guidance (CFG) routinely fails to synthesize unobserved concept combinations, exact decomposed scoring generalizes robustly at the cost of severe computational overhead. In this work, we reveal that compositional binding is not a uniform process but a highly localized phase transition. We identify the semantic bifurcation window - the precise temporal interval where joint and decomposed vector fields meaningfully diverge. Exploiting this dynamic, we propose surgical guidance, a hybrid sampling strategy that restricts exact multi-pass scoring strictly to this critical window. On an OOD bi-digit MNIST testbed, surgical guidance achieves state-of-the-art compositional fidelity at a fraction of the inference cost, yielding a +5.3% absolute improvement in pairwise accuracy over the joint baseline by intervening during just the first 15% of the diffusion trajectory. Furthermore, our empirical analysis uncovers a fundamental topological divide: diffusion models (SDEs) force conceptual resolution immediately at peak noise, whereas Conditional Flow Matching (ODEs) delays structural binding until intermediate features emerge, establishing a new temporal framework for accelerating large-scale generative decoding.
Figure 2: Divergence between joint and decomposed scoring vector fields.
Figure 3: Visualizing attractor collapse during the semantic bifurcation window.
Method
(2,8)
(3,2)
(5,3)
(8,5)
(2,2)
(3,3)
(5,5)
(8,8)
Total
CFDG (joint)
74
47
45
57
32
12
50
6
323
CFDG (decomposed)
27
28
29
30
15
3
37
2
171
CFM (joint)
19
23
28
21
14
14
18
21
158
CFM (decomposed)
18
21
9
14
8
2
15
12
99
Table 1: Out-of-distribution compositional generalization errors. Values represent the number of misclassified samples out of 300 novel held-out pairs (lower is better).
Figure 4: Surgical guidance mechanism across CFM and DDIM frameworks.
Method
(2,8)
(3,2)
(5,3)
(8,5)
(2,2)
(3,3)
(5,5)
(8,8)
Mean
CFDG (surgical)
16.9
14.5
13.5
13.9
10.5
18.7
11.8
23.5
15.4
CFM (surgical)
16.9
15.7
15.5
16.4
12.7
19.0
14.3
19.6
16.3
Table 2: Fréchet Inception Distance (FID) for surgical sampling on out-of-distribution pairs. Lower values indicate stronger distributional alignment with the training set.
The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to produce samples from compositionally-defined target distributions such as a geometric combination of the source distributions. In this work, we argue that this task is often infeasible for vanilla conditional diffusion models: we conjecture that no inference-time technique can efficiently produce samples from the target distribution in certain well-motivated settings. This idea is supported by theory-guided generalization arguments and carefully-designed experiments on both synthetic and realistic data. In particular, while recent methods such as Feynman-Kac correction reduce inference-time approximation error, our results show that score estimation error has a more catastrophic effect on performance when the target distribution is out-of-distribution with respect to the sources, highlighting the need for a different approach to this task.
Duncan Soiffer, Chandler Squires, Yuan Guan +2
Machine Learning Department, Carnegie Mellon University · Valence Labs · Department of Computer Science, University of Manchester
Compositional generalization requires models to produce novel configurations from familiar parts. In diffusion models, prior compositional generation methods typically assume that the relevant concepts or conditioning signals are already available. We instead ask whether a pretrained diffusion model can discover query-specific concepts from the time-indexed scores it learns for the noisy marginals pt(xt) and compose them at test time. Given a single out-of-distribution query, our method performs gradient ascent on sθ(xt,t)≈∇xtlogpt(xt) at multiple noising timesteps to recover local density modes, maps these modes into clean-space Gaussians, greedily selects relevant prototypes with a submodular likelihood objective, and combines them into a product-of-experts (PoE) teacher model with an analytic score. This teacher model can be sampled directly through classifier-free guidance or used to generate a sample pool for training a new class embedding and low-rank adapter. On held-out composition benchmarks built from ColorMNIST and CelebA, both the analytic PoE sampler and the low-rank adapted model outperform query-only and nearest trained-class baselines. These results suggest that the time-indexed score geometry of the diffusion model contains reusable density-mode concepts that support test-time compositional generation without a predefined concept library.
Zekun Wang, Anant Gupta, Tianyi Zhu +1
Georgia Institute of Technology · University of Virginia
Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of which design choices critically determine compositional generalization in image and video generation. By isolating independent design axes, we identify two key factors strongly associated with compositional success: (i) whether the training objective operates on a discrete or continuous distribution, and (ii) the completeness of conditioning information about constituent factors during training. We also show that relaxing the discrete loss with an auxiliary continuous latent objective can partially recover compositional performance in discrete models like MaskGIT. Our findings, corroborated by diverse compositional tasks and preliminary evidence in world models and LLMs, motivate a shift toward continuous objectives for compositional generalization.
Karim Farid, Rajat Sahay, Yumna Ali Alnaggar +4
University of Freiburg · Bosch Center for Artificial Intelligence · Inria, École Normale Supérieure, CNRS, PSL Research University