ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
Authors: Arshia Hemmat, Amirhossein Vahidi, Amitis Shidani, Mohammad Vali Sanian, Hesam Asadollahzadeh, Aryan Yazdan Parast, Mohammad Lotfollahi
Organizations: Wellcome Sanger Institute · Cambridge Stem Cell Institute, University of Cambridge · Cambridge Centre for AI in Medicine, University of Cambridge · University of Oxford · Apple · Department of Computer Science, University of Helsinki · Institute for Molecular Medicine Finland (FIMM), University of Helsinki · School of Computing and Information Systems, University of Melbourne · Department of Medicine, University of Cambridge
Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.
Figures & tables
Figure 1 : ORCA aligns the diffusion latent with a low-rank visual target via a text-conditioned projector K(Δy) built from the T5 − CLIP residual. On DiT-L/2, ORCA reaches the vanilla 400K FID and GenEval in 4.8× and 3.6× fewer steps, with zero inference overhead.
Figure 2 : (a) Mean principal angle (in degrees) between the subspaces induced by paired prompts that differ along a controlled axis. Pairs are grouped by perturbation type: color attribute swap, object change, relation addition, and multi-axis change. The angles increase systematically with the semantic distance between prompts, with multi-axis changes producing the largest deviations and color swaps the smallest. The QR map is therefore input-dependent in a structured way: linguistically related prompts induce nearby readout subspaces, while linguistically distant prompts induce more separated ones. (b) Geometric illustration of the orthogonal decomposition of Equation ( 9 ) with rank r=64 and d=768 .
DiT-B/2 (130M)
DiT-L/2 (458M)
U-ViT-L (287M)
Setting
Iter.
FID ↓
GenEval ↑
Iter.
FID ↓
GenEval ↑
Iter.
FID ↓
GenEval ↑
Vanilla diffusion training
Vanilla
100K
35.49
0.123
100K
27.77
0.162
100K
38.72
0.130
Vanilla
150K
33.55
0.145
150K
24.26
0.191
150K
32.48
0.148
Vanilla
200K
29.95
0.167
200K
24.71
0.195
200K
30.95
0.161
Vanilla
400K
29.15
0.174
400K
24.01
0.247
400K
29.91
0.189
Table 1 : Main results on MS-COCO 256×256 . Results are grouped by diffusion backbone. For each backbone, we report the training checkpoint, FID-30K, and GenEval. † ORCA adds zero parameters and FLOPs at inference. Best per backbone in bold. REPA REG ORCA (ours)
Method
Single
Two
Count
Colors
Position
Color attr.
Overall
Vanilla
0.516
0.038
0.116
0.431
0.025
0.023
0.191
REPA
0.638
0.040
0.116
0.620
0.050
0.030
0.249
ORCA ( r=16 )
0.622
0.043
0.119
0.588
0.043
0.038
0.242
ORCA ( r=32 )
0.613
0.083
0.181
0.582
0.050
0.068
0.263
ORCA ( r=64 )
0.634
0.061
0.159
0.638
0.073
0.068
0.272
ORCA ( r=128 )
0.669
0.068
0.194
0.641
0.058
0.048
0.279
Table 2 : GenEval per-task breakdown on DiT-L/2 at 150K steps for all methods (matched to the rank-ablation budget). Gains concentrate on tasks that require compositional reasoning ( Two objects , Counting , Position , Color attribution ); single-object accuracy is comparatively saturated and improves only modestly. We additionally report ORCA at multiple ranks r to show the gain is not idiosyncratic to one rank choice. Best per column in bold.
Figure 3 : ORCA design-choice sensitivity (FID-30K). Each panel sweeps a single hyperparameter while the other two are held at their default; stars mark the default configuration adopted in the main results. Left: the auxiliary-loss weight λ exhibits a clean optimum at λ=1.0 , with both smaller and larger values degrading FID. Centre: the bottleneck rank r is non-monotonic, with r=64 as a sharp minimum. Right: applying the alignment loss at intermediate depth (block 8 of 24 ) is strongly preferred; alignment at the final block collapses the model entirely (FID 64.16 ). Full per-metric tables are in Appendix D .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Specification
Parameters
Textual projection W
Linear( dT→dC )
dT⋅dC+dC
Basis MLP gϕ , hidden
Linear( dC→hg ) + LayerNorm + GELU
dChg+hg+2hg
Basis MLP gϕ , output
Linear( hg→d⋅n ), Xavier init (gain 0.01 )
hg⋅d⋅n+d⋅n
Orthonormalisation
QR ( torch.linalg.qr ) on d×n matrix
0 (deterministic)
PCA projection Pn , vˉ
Precomputed from frozen DINOv2
0 (no gradient)
Appendix
Table 3 : ORCA module architecture. The textual projection W is a single linear layer mapping the T5 caption embedding into the dC -dimensional CLIP space. The basis MLP gϕ is a 2-layer MLP with LayerNorm and GELU; its output layer is initialised with Xavier gain 0.01 and zero bias, producing a small pre-QR candidate matrix at initialisation. The QR step and the PCA projection contribute no learnable parameters.
Hyperparameter
Value
Notes
Diffusion training (inherited)
Optimiser
AdamW
β1=0.9 , β2=0.999
Learning rate
1×10−4
cosine schedule with linear warmup
Warmup steps
5,000
Batch size (global)
256
Total training iterations
400,000
for the baselines result
Appendix
Table 4 : Hyperparameters for ORCA training. The auxiliary-loss weight λ and rank n are the two hyperparameters specific to ORCA; all others are inherited from the diffusion backbone.
Model
Iter.
Single
Two
Count
Colors
Position
Color attr.
Overall
DiT-B/2
DiT-B/2
100K
0.381
0.023
0.072
0.250
0.008
0.005
0.123
DiT-B/2
150K
0.397
0.033
0.097
0.301
0.020
0.020
0.145
DiT-B/2
200K
0.475
0.030
0.100
0.356
0.020
0.020
0.167
DiT-B/2
400K
0.491
0.020
0.106
0.378
0.028
0.020
0.174
+ REPA
250K
0.559
0.058
0.097
0.497
0.038
0.038
0.214
Appendix
Table 5 : Per-task GenEval breakdown for the main configurations in Table 1 . “Single” = single object, “Two” = two objects, “Count” = counting, “Colors” = single-object color, “Position”, “Color attr.” = color attribution.
Rank r
FID ↓
GenEval ↑
CLIP-L/14 ↑
PickScore ↑
16
18.79
0.242
0.1076
17.16
32
19.33
0.263
0.1063
17.12
64
16.76
0.272
0.1042
17.11
128
18.29
0.279
0.1045
17.14
Appendix
Table 6 : Effect of subspace rank r on DiT-L/2 at 150K steps. FID is best at r=64 , while GenEval is highest at r=128 . Best per column in bold.
Layer
FID ↓
GenEval ↑
CLIP-B/32 ↑
PickScore ↑
1 (very early)
36.63
0.130
0.2159
17.33
8 (default)
27.61
0.165
0.2172
17.25
16 (middle)
30.23
0.162
0.2165
17.23
24 (last)
64.16
0.054
0.2156
17.40
Appendix
Table 7 : Layer placement of the auxiliary loss on DiT-L/2 at 50K training steps. Best per column in bold.
λ
FID ↓
GenEval ↑
CLIP ↑
PickScore ↑
0.1
32.52
0.135
0.2163
17.29
0.5
28.43
0.182
0.2164
17.26
1.0
27.61
0.165
0.2172
17.25
2.0
31.42
0.172
0.2151
17.22
5.0
31.21
0.157
0.2159
17.25
Appendix
Table 8 : Effect of the auxiliary-loss weight λ on DiT-L/2 at 50K steps. The λ=1.0 configuration is the default used for the ORCA runs reported in Table 1. Best per column in bold.
Target
Loss
FID ↓
GenEval ↑
Learnable linear (no reg.)
MSE
collapsed
collapsed
Learnable linear + VICReg
MSE
29.49
0.145
Frozen PCA (ours)
MSE
16.76
0.272
Frozen PCA
Gaussian NLL
23.31
0.257
Appendix
Table 9 : Loss function and target construction on DiT-L/2. Frozen PCA with MSE is the strongest configuration. Best per column in bold.
CLIP
T5
DINOv2
FID ↓
GenEval ↑
ViT-B/32
T5-XL
ViT-B/14
16.76
0.272
Appendix
Table 10 : Preliminary sensitivity check under an additional pretrained-encoder configuration on DiT-L/2.
Figure 4 : Projection metrics over training. Energy ratio (orange), aligned-component cosine (blue), and residual-target cosine (red, dashed). The aligned component captures 80% of ∥ht∥2 in a rank- 64 subspace and reaches cosine 0.80 with the DINOv2 PCA target. The residual-target cosine is zero by construction and is included only as an implementation sanity check; it should not be interpreted as evidence that the orthogonal complement contains no visual information.
Model
Iter.
FID ↓
GenEval ↑
DiT-L/2
DiT-L/2
100K
27.77±0.16
0.162±0.008
DiT-L/2
150K
24.26±0.17
0.191±0.008
DiT-L/2
200K
24.71±0.16
0.195±0.009
DiT-L/2
400K
24.01±0.16
0.247±0.009
+ REPA
400K
20.05±0.12
0.275±0.009
Appendix
Table 11 : Bootstrap error bars on FID and GenEval for the DiT-L/2 configurations of Table 1 . FID standard deviations come from 100 bootstrap resamples of the 30K generated feature pool; GenEval standard deviations come from 1,000 bootstrap resamples of the per-image correctness scores. All standard deviations are 1 - σ , unbiased ( ddof=1 ). DiT-B/2 and U-ViT-L bootstrap evaluations are pending and will be added if the runs complete in time.
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary-condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify-then-Diffuse (RTD), a training-free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft-Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout-agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state-of-the-art compositional fidelity and robust gains. On the AE-Bench object pair subset, RTD improves BLIP-VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3× faster. Code will be released at https://github.com/Z-yiwei/rectify-then-diffuse
Ning Zhu, An Chen, Mengfei Zhao +4
Glasgow College, University of Electronic Science and Technology of China · School of Mathematical Sciences, University of Electronic Science and Technology of China
Text-to-image (T2I) diffusion models often fail to faithfully render explicit textual descriptions, instead defaulting to strongly learned visual priors due to a phenomenon referred to as concept association bias. We show that such bias is particularly strong for one-and-only (OAO) objects, entities that exist in a single canonical form, such as celestial bodies, landmarks, and artworks. The deeply ingrained visual identity for these concepts often resists modification through prompting alone. Addressing this challenge, we first identify through an information-theoretic analysis that the final text embedding discards concept-level information present in the intermediate-layer text representations, reducing the mutual information available to the subsequent denoising process. We then propose Intermediate Text Representation (IR)-guided diffusion, which injects intermediate hidden states of the text encoder into the conditioning signal during early denoising steps, recovering suppressed concepts without any additional training, optimization, or external models. To systematically evaluate the challenging task of aligning generative outputs with unusual prompts for OAO objects, we introduce OAO-AttackBench, a benchmark comprising counterfactual prompts that directly conflict with the core visual identity of OAO objects. Experiments on four benchmarks, including OAO-AttackBench, show that our method achieves up to a 19.1 percentage-point improvement in VQAScore while preserving generation fidelity and human preference. Project page: https://soyoun-won.github.io/one-and-only-ir-guidance/.
Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relations between entities, attributes, actions, and motion directions. We hypothesize that these failures need not be addressed by retraining the generator, but can instead be mitigated by steering the denoising process using the model's own internal grounding signals. We propose \textbf{CVG}, an inference-time guidance method for improving compositional faithfulness in frozen text-to-video models. Our key observation is that cross-attention maps already encode how prompt concepts are grounded across space and time. We train a lightweight compositional classifier on these attention features and use its gradients during early denoising steps to steer the latent trajectory toward the desired composition. Built on a frozen VLM backbone, the classifier transfers across semantically related composition labels rather than relying only on narrow category-specific features. CVG improves compositional generation without modifying the model architecture, fine-tuning the generator, or requiring layouts, boxes, or other user-supplied controls. Experiments on compositional text-to-video benchmarks show improved prompt faithfulness while preserving the visual quality of the underlying generator.
Ariel Shaulov, Eitan Shaar, Amit Edenzon +2
1Tel-Aviv University · 2Independent Researcher · 3Bar Ilan University +1