Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of which design choices critically determine compositional generalization in image and video generation. By isolating independent design axes, we identify two key factors strongly associated with compositional success: (i) whether the training objective operates on a discrete or continuous distribution, and (ii) the completeness of conditioning information about constituent factors during training. We also show that relaxing the discrete loss with an auxiliary continuous latent objective can partially recover compositional performance in discrete models like MaskGIT. Our findings, corroborated by diverse compositional tasks and preliminary evidence in world models and LLMs, motivate a shift toward continuous objectives for compositional generalization.
Figures & tables
Figure 1 : Compositional Generalization Analysis. We compare discrete masking (MaskGIT) and continuous diffusion (DiT) on novel compositions of three binary CelebA factors: gender, hair color, and smile. Models are trained on four observed combinations ( blue ) and evaluated on held-out Level-1 ( pink , one-factor change) and Level-2 ( red , two-factor change) compositions. DiT (row 1) generalizes compositionally, while MaskGIT (row 2) struggles.
Model
Masking-based
Training loss
Latent Distribution
MaskGIT
✓
Categorical NLL
discrete
GIVT
✓
GMM-NLL
continuous
MAR
✓
Diffusion
continuous
DiT
✗
Diffusion
continuous
Table 1: Summary of generative models.
Figure 2 : DiT exhibits compositional generalization regardless of the type of tokenizer used. While the training dynamics differ, DiT shows compositional generalization at the end of training. The blue , pink , and red curves show linear probe accuracies for the training data, level-1 compositions, or level-2 compositions, respectively. Results are consistent for MAR ( Figure 7 ) and video ( Figure 8 ).
Figure 3 : Compositional generalization performance on Shapes2D across different model architectures. Models that learn continuous distributions (DiT, MAR, and GIVT) consistently show better level-2 compositions than MaskGIT, with the decisive shift in performance occurring at the categorical-to-continuous intervention. The blue , pink , and red curves denote training, level-1, and level-2 compositions, respectively. Consistent results are observed for CLEVRER-Kubric ( Figure 9 ).
Figure 4 : Comparison of conditioning information levels and their impact on compositional generalization in DiT on Shapes2D. (a) Continuous (full-information) conditioning leads to uniform convergence across all compositions. (b) Label dropout conditioning leads to inconsistent generalization; several unseen compositions fail. (c) Discrete (quantized) conditioning leads to partial generalization, with some failing samples. (d) Discrete (quantized) conditioning with dropout, the most severe loss of information, leads to the most failure. Shaded areas indicate standard deviation across three different seeds. We provide additional results in Section C.5 . The blue curves show performance on training data, pink curves depict level-1 compositions, and red curve denotes level-2.
Figure 5 : An overview of MaskGIT combined with the JEPA-based training objective. We apply the JEPA loss at specific layers ( l ) on an intermediate masked token representation in the transformer (HC(l)) and train a lightweight predictor to reconstruct target states (HT(l)) using MSE as an error metric and a stop-gradient signal to avoid representation collapse.
Figure 6 : Comparison of linear probe accuracy for MaskGIT with ( 6(a) ) Standard and ( 6(b) ) JEPA-based training objectives for Shapes2D. The JEPA-based objective clearly enhances compositional abilities even though it cannot fully compensate the problems introduced by the discrete space. We also see lower polysemanticity between attention heads ( 6(c) ) and a decreasing Jaccard trend over shared circuits ( 6(d) ). The blue curves show performance on training data, pink curves depict level-1 compositions, and red curve denotes level-2 compositions. Consistent videos results shown in Figure 13 .
Appendix figures & tables32 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : MAR exhibits generalization on both level-1 and level-2 compositions in Shapes2D regardless of whether (a) discrete, (b) continuous, or (c) no tokenizer is used. The blue curves show performance on training data, pink curves depict level-1 compositions, and red curve denotes level-2 compositions.
Figure 8 : DiT exhibits generalization on both level-1 and level-2 compositions in CLEVRER-Kubric regardless of whether (a) discrete, (b) continuous, or (c) no tokenizer is used. The blue curves show performance on training data, pink curves depict level-1 compositions, and red curve denotes level-2 compositions.
Figure 9 : Comparison of linear probe accuracy for CLEVRER-Kubric across ( 9(a) ) MaskGIT, ( 9(b) ) DiT, and ( 9(c) ) MAR. Models leveraging a continuous latent space (DiT, and MAR) show better level-2 compositions than MaskGIT. The blue curves show performance on training data, pink curves depict level-1 compositions, and red curve denotes level-2 compositions.
Figure 10 : Compositional generalization across different training conditions for MaskGIT on Shapes2D. We see that MaskGIT struggles to compositionally generalize despite varying multiple different training parameters.
Model
Seen Mean
Level-1 Novel Mean
Level-2 Novel
(011, 110, 101)
(111)
AR (discrete)
86.7
38.5
6.2
AR-GMM (continuous)
90.2
64.1
46.9
Appendix
Table 2 : Probe accuracy (%) of causal autoregressive models with a discrete (categorical) and a continuous (Gaussian mixture) objective on Shapes2D.
Unmasking
000
001
010
100
Seen Mean
011
101
110
111
Novel Mean
Confidence
87.8
94.3
82.9
77.3
85.6
64.9
54.0
38.4
0.7
39.5
Random order
87.8
95.0
86.1
74.3
85.8
45.7
41.2
26.3
2.0
28.8
Appendix
Table 3 : Probe accuracy (%) of MaskGIT on Shapes2D with confidence-based vs. random unmasking order (256 samples per composition).
Figure 11 : Comparison of linear probe accuracy for DiT trained on CLEVRER-Kubric with ( 11(a) ) Explicit conditioning and ( 11(b) ) Implicit conditioning mechanisms. Training with explicit label information allows for better level-2 compositions. The blue curves show performance on training data, pink curves depict level-1 compositions, and red curve denotes level-2 compositions.
Figure 12 : Qualitative Results of the concept dropout with DIT on Shapes2D. The visualization demonstrates that when contextual information about all concepts is incomplete, the model struggles to generalize compositionally. A typical failure mode is the substitution of novel concept combinations with the closest seen ones—for instance, the 110 configuration, which should correspond to a novel blue big triangle, is instead rendered as a red big triangle, likely due to its proximity in the training distribution. Similar behavior is happening to the 101, which should correspond to a novel red small triangle, sometimes result in a red small circle or red big triangle.
Model
Conditioning
Seen
Novel L1
Novel L2
All Novel
DiT
Factors
97.3
99.0
96.9
98.4
Language
74.2
71.4
48.4
65.6
Language (partial descriptions)
76.6
62.5
1.6
47.3
MaskGIT
Factors
85.6
52.4
0.7
39.5
Language
84.1
52.0
1.6
39.4
Language (partial descriptions)
85.2
19.9
0.3
15.0
Appendix
Table 4 : Probe accuracy (%) on Shapes2D with explicit factor conditioning vs. CLIP text conditioning. L1/L2 denote Level-1/Level-2 novel compositions; All Novel averages over all unseen compositions.
Figure 13 : Comparison of linear probe accuracy for MaskGIT with Standard ( 13(a) ) and JEPA-based ( 13(b) ) training objectives for CLEVRER-Kubric. The JEPA-based training objective clearly enhances compositional abilities, and lowers polysemanticity ( 13(c) ) even though it cannot fully compensate the problems introduced by the discrete space. The blue curves show performance on training data, pink curves depict level-1 compositions, and red curve denotes level-2 compositions.
Layers
Accuracy
{6}
27.61
{8}
26.35
{6,8}
38.16
{7,9}
36.62
{6,8,10}
53.52
{7,9,11}
56.27
Appendix
Table 5 : Maximum linear probe accuracy after applying the JEPA loss on different layers and combinations of layers for a MaskGIT model trained on CLEVRER-Kubric
Figure 14 : Effect of weighting factor on probe accuracy
Figure 15 : Visualization of MaskGIT circuits for a large red cube and small blue sphere trained with the standard training objective versus the JEPA-based training objective.
Color-ablated
Shape-ablated
MaskGIT+JEPA
Color-varying
↓38 pp
↓ 6pp
Shape-varying
↓ 5pp
↓41 pp
Standard MaskGIT
Color-varying
↓ 18pp
↓ 14pp
Shape-varying
↓ 15pp
↓ 17pp
Appendix
Table 6 : Level-2 probe accuracy drops after ablating factor-specific circuits. Bold entries mark factor-aligned ablations. MaskGIT+JEPA shows selective degradation; standard MaskGIT shows uniform drops, consistent with entangled representations.
Corruption
Strength
Shape
Color
Size
Joint
Blur ( σ )
1
98.9
100.0
89.5
88.4
2
99.2
100.0
91.7
90.9
4
73.1
100.0
89.3
64.7
8
0.0
100.0
50.0
0.0
Noise ( σ )
0.1
97.4
100.0
93.9
91.2
0.25
91.1
100.0
78.2
75.5
Appendix
Table 7 : Probe accuracy (%) on real Shapes2D test images under increasing corruptions. Joint denotes all three factors being classified correctly.
Model
(0,0,0)
(0,0,1)
(0,1,0)
(1,0,0)
(1,0,1)
(1,1,0)
(0,1,1)
(1,1,1)
DiT
100%
100%
100%
100%
100%
100%
100%
100%
MaskGIT
94.7%
97%
94.7%
94%
79.49%
62.6%
58.85%
30%
Appendix
Table 8 : Performance comparison of DiT and MG across different compositions.
red, orange, yellow, lime green, green, cyan, blue
indigo, purple, pink
Shape
cube, cylinder, sphere
cube, cylinder
sphere
Size
0–7 (ordinal)
0, 1, 2, 3, 4, 5
6, 7
Appendix
Table 11 : Factorization of the data into Super Groups . This partitioning is used to evaluate compositional generalization across disjoint super groups.
Figure 17 : Shapes3D results of DiT. Dark blue : all factors from Supergroup 0; light blue : one factor from Supergroup 1; pale orange : level-1 compositions; and dark red : level-2 novel compositions.
Figure 18 : Shapes3D results of DiT. Dark blue : all factors from Supergroup 0; light blue : one factor from Supergroup 1; pale orange : level-1 compositions; and dark red : level-2 novel compositions.
Figure 19 : Shapes3D results of DiT. Dark blue : all factors from Supergroup 0; light blue : one factor from Supergroup 1; pale orange : level-1 compositions; and dark red : level-2 novel compositions.
Figure 20 : Shapes3D results of MaskGIT. Dark blue : all factors from Supergroup 0; light blue : one factor from Supergroup 1; pale orange : level-1 compositions; and dark red : level-2 novel compositions.
Figure 21 : Shapes3D results of MaskGIT. Dark blue : all factors from Supergroup 0; light blue : one factor from Supergroup 1; pale orange : level-1 compositions; and dark red : level-2 novel compositions.
Figure 22 : Shapes3D results of MaskGIT. Dark blue : all factors from Supergroup 0; light blue : one factor from Supergroup 1; pale orange : level-1 compositions; and dark red : level-2 novel compositions.
Figure 23 : Qualitative Results on CelebA
Real train
Real novel
Gen. novel
Model
\astrosun ←
\astrosun ↑
\fullmoon →
\fullmoon ↑
\astrosun →
\fullmoon ←
\astrosun →
DiT
0.190
0.190
0.060
0.060
0.470
0.030
MaskGIT
0.380
0.260
0.110
0.050
0.180
0.020
\fullmoon ←
DiT
0.130
0.020
0.340
0.040
0.040
0.430
MaskGIT
0.150
0.030
0.560
0.100
0.020
0.140
Appendix
Table 12 : Nearest-neighbor compositional retrieval metric (CRA) comparing generated videos to real videos. Factors are time of day: Day \astrosun , Night \fullmoon ; directions: Left ← , Straight ↑ , Right → . Rows show a generated novel split evaluated against real train and real novel classes. The Hit@1 (correct target novel class) is bolded; black bold indicates better results.
Method
FVD (Seen) ↓
FVD (Unseen) ↓
DiT
730.30
2257.51
MG
2320.95
3197.74
Appendix
Table 13 : FVD on seen and unseen splits on CoVLA.
Figure 24 : Qualitative results with Orbis on CoVLA. Two novel compositions are shown, \astrosun → and \fullmoon ← , each conditioned on a 7-frame initial context and predicting 5 frames ahead. Orbis-DiT follows the target compositions, whereas Orbis-MaskGIT often struggles.
Figure 25 : Example system prompt, unsuccessful model output, and verifier feedback for our language task with compositional rules: target restriction and red-suit value doubling.
Figure 26 : Example system prompt, successful model output, and verifier feedback for a single-rule training instance of our language task: target restriction without red-suit value doubling.
The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to produce samples from compositionally-defined target distributions such as a geometric combination of the source distributions. In this work, we argue that this task is often infeasible for vanilla conditional diffusion models: we conjecture that no inference-time technique can efficiently produce samples from the target distribution in certain well-motivated settings. This idea is supported by theory-guided generalization arguments and carefully-designed experiments on both synthetic and realistic data. In particular, while recent methods such as Feynman-Kac correction reduce inference-time approximation error, our results show that score estimation error has a more catastrophic effect on performance when the target distribution is out-of-distribution with respect to the sources, highlighting the need for a different approach to this task.
Duncan Soiffer, Chandler Squires, Yuan Guan +2
Machine Learning Department, Carnegie Mellon University · Valence Labs · Department of Computer Science, University of Manchester
Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relations between entities, attributes, actions, and motion directions. We hypothesize that these failures need not be addressed by retraining the generator, but can instead be mitigated by steering the denoising process using the model's own internal grounding signals. We propose \textbf{CVG}, an inference-time guidance method for improving compositional faithfulness in frozen text-to-video models. Our key observation is that cross-attention maps already encode how prompt concepts are grounded across space and time. We train a lightweight compositional classifier on these attention features and use its gradients during early denoising steps to steer the latent trajectory toward the desired composition. Built on a frozen VLM backbone, the classifier transfers across semantically related composition labels rather than relying only on narrow category-specific features. CVG improves compositional generation without modifying the model architecture, fine-tuning the generator, or requiring layouts, boxes, or other user-supplied controls. Experiments on compositional text-to-video benchmarks show improved prompt faithfulness while preserving the visual quality of the underlying generator.
Ariel Shaulov, Eitan Shaar, Amit Edenzon +2
1Tel-Aviv University · 2Independent Researcher · 3Bar Ilan University +1
Systematic out-of-distribution (OOD) generation remains a critical bottleneck for continuous-time generative models. While standard joint classifier-free guidance (CFG) routinely fails to synthesize unobserved concept combinations, exact decomposed scoring generalizes robustly at the cost of severe computational overhead. In this work, we reveal that compositional binding is not a uniform process but a highly localized phase transition. We identify the semantic bifurcation window - the precise temporal interval where joint and decomposed vector fields meaningfully diverge. Exploiting this dynamic, we propose surgical guidance, a hybrid sampling strategy that restricts exact multi-pass scoring strictly to this critical window. On an OOD bi-digit MNIST testbed, surgical guidance achieves state-of-the-art compositional fidelity at a fraction of the inference cost, yielding a +5.3% absolute improvement in pairwise accuracy over the joint baseline by intervening during just the first 15% of the diffusion trajectory. Furthermore, our empirical analysis uncovers a fundamental topological divide: diffusion models (SDEs) force conceptual resolution immediately at peak noise, whereas Conditional Flow Matching (ODEs) delays structural binding until intermediate features emerge, establishing a new temporal framework for accelerating large-scale generative decoding.
Nguyen-Thanh-Luong Doan, Quang-Vu Nguyen, Tang-Phu-Quy Le +1
The University of Danang - Vietnam-Korea University of Information and Communication Technology