TILT: Model-Intrinsic Reward Alignment For Compositional Diffusion
Authors: Debottam Dutta, Jianchong Chen, Jaehoon Hahm, Romit Roy Choudhury
Organizations: Electrical and Computer Engineering, University of Illinois Urbana-Champaign, Urbana, IL, USA · Zhejiang University, China · Physics, University of Illinois Urbana-Champaign, Urbana, IL, USA
Consider conditional generation p(x∣C={c1,c2,…ck}) where C is a prompt composed of multiple concepts ci. Diffusion models often struggle with compositional prompts, producing samples in which some concepts dominate while others are missing or weakly represented. Prior work attributes these failures to mode collision, where single-concept modes of p(x∣ci) overlap with modes of the joint p(x∣C). To seek out collision-free modes of p(x∣C), or "pure modes", corrector-based approaches have attempted to suppress collisions at intermediate diffusion times. However, local corrections are often heuristic and do not necessarily steer the generation to a "pure mode" in the final data space. Derived from a principled formulation, we present TILT (Test-time model-Intrinsic reward aLignment via Tilting), a training-free framework that poses eventual pure mode sampling as a reward for intermediate-time alignment. This reward offers valuable advantages: (1) it is intrinsic to the model, hence external reward models need not be trained by modality-specific datasets, (2) it yields a closed-form target under a variational approximation, which makes it realizable through standard diffusion sampling, and (3) it is interpretable, hence amenable to preference-based modifications. Project page: https://debottam-dutta7.github.io/tilt_web/
Figures & tables
Figure 1: The correct intermediate-time distribution. Left: in data space ( t=0 ), the joint p0(⋅∣C) (blue) shares modes with individual concept distributions p0(⋅∣c1) (orange) and p0(⋅∣c2) (brown). The distribution p~0 (purple) describes the pure modes – the modes of the joint that no single concept explains. Right: at t=T/3 , dark-purple is the diffusion pushforward of p~0 ( divide, then diffuse ), while green is the corrector distribution from Eq. 1 , which follows diffuse, then divide . The two distributions diverge – a corrector defined at t>0 therefore steers the sampler towards some other distribution, not the noised “pure mode” target.
Figure 2: Density plot of DINOv2 scores between each sub-prompt c1,c2 and the generated image using SDXL. This visualizes the failure in compositional generation.
Method
T2V ↑
HPSv3 ↑
CFG
0.3934
5.5029
CFG++
0.5073
5.9806
R2F
0.4799
5.8176
SuperDiff
0.1589
0.082
CO3
0.5014
4.3026
DOS
0.5025
5.5744
Table 1: Quantitative results on GenEval prompts.
Color
Shape
Texture
Complex
Method
T2V ↑
HPSv3 ↑
T2V ↑
HPSv3 ↑
T2V ↑
HPSv3 ↑
T2V ↑
HPSv3 ↑
CFG
0.5183
5.1216
0.3565
6.1377
0.5529
6.5021
0.6238
7.5702
CFG++
0.5668
5.7358
0.3839
6.9679
0.5839
7.1200
0.6307
8.1785
R2F
0.5328
5.5146
0.3452
6.5545
0.5573
6.7431
0.6234
8.0638
CO3
0.5938
4.4736
0.3895
6.1982
0.4252
6.2801
0.6313
6.6193
DOS
0.5028
4.8979
0.3865
6.0388
0.6199
6.4046
0.6603
6.4046
Table 2: Quantitative results on T2I-CompBench prompts.
Figure 3: Qualitative comparison on GenEval and T2I-CompBench. The top two rows use prompts from GenEval and the bottom four rows use prompts from T2I-CompBench (shape, texture, color, and complex subsets); each row’s prompt is shown on the left. All methods use the SDXL backbone and share the same random seed within each row.
Method
BLIP-VQA ↑
T2V ↑
HPSv3 ↑
SDXL
0.339
0.483
5.50
FKS-CLIP
0.450
0.563
5.89
FKS- TILT
0.446
0.530
5.91
Table 3: Comparison with extrinsic-reward steering on GenEval.
Figure 7
Method
CLAP ↑
FAD ↓
KL ↓
FD ↓
AudioLDM
0.4508
4.9867
0.7782
3.0924
AudioLDM- TILT
0.5350
3.6415
0.6213
2.7377
AudioLDM- TILT (best of 8)
0.6110
3.0848
0.6176
1.9624
SCORE- TILT (8 particles)
0.5205
3.8210
0.5890
1.8640
SCORE-CLAP (8 particles)
0.6258
3.1874
0.5852
1.7321
Table 4: Quantitative results on audio generation using prompts from AudioCaps.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Additional success cases of TILT . The top two rows use prompts from GenEval (two objects and color attribute) and the bottom seven rows use prompts from T2I-CompBench (texture and complex subsets); each row’s prompt is shown on the left. All methods use the SDXL backbone and share the same random seed within each row.
Figure 7: Additional success cases of FKS with our reward on GenEval. Each row’s prompt is shown on the left. With the CLIP reward, FKS keeps samples that miss or merge an object (e.g., the person with an apple for a head, or the spoon in place of the knife) or bind a color to the wrong object (e.g., the green umbrella), while our reward keeps samples that render every object with its intended color.
Figure 8: Failure cases of TILT . The top three rows use prompts from GenEval (color attribute, position, and two objects) and the bottom four rows use prompts from T2I-CompBench (color, shape, texture, and complex subsets); each row’s prompt is shown on the left. TILT fails on these prompts for all three evaluation seeds. Typical failure modes include replacing one concept with a copy of the other, dropping a concept entirely, fusing two concepts into a hybrid and swapping attributes.
Figure 9: Failure cases of FKS ( TILT ) on GenEval. Each row’s prompt is shown on the left. FKS with our reward fails on these prompts for both evaluation seeds. The top three rows are color-attribute prompts, where the selected sample binds a color to the wrong object (e.g., the swapped colors of the cake and the chair); the bottom three rows are position prompts, where the selected sample drops an object (the parking meter) or places the objects in the wrong relation (the cell phone on the chair and the tie on the bat).
Figure 10: Additional qualitative comparison on AudioCaps. Mel spectrograms of samples from SCORE ( Jung et al., 2025 ) steered by the CLAP reward (SCORE-CLAP) and by our reward (SCORE- TILT ), both on AudioLDM; each row’s prompt is shown on the left.
energy_num_samples
BLIP-VQA ↑
T2V ↑
HPSv3 ↑
Base: default
2
0.3942
0.5192
6.122
4
0.3900 ( − 0.0041)
0.4952 ( − 0.0240 † )
6.049 ( − 0.072)
Base: default with energy_adaptive_weights , energy_adaptive_zeta=3
2
0.3898
0.5178
6.120
4
0.3890 ( − 0.0009)
0.5067 ( − 0.0111)
6.069 ( − 0.051)
Appendix
Table 5: Effect of energy_num_samples , the number of (τ,ϵ) draws used to estimate the energy reward, on GenEval. As the number of energy samples increases, the variance of the estimate decreases, leading to more stable performance.
energy_adaptive_zeta
BLIP-VQA ↑
T2V ↑
HPSv3 ↑
Base: default
off
0.3942
0.5192
6.122
3
0.3898 ( − 0.0043)
0.5178 ( − 0.0014)
6.120 ( − 0.002)
5
0.3932 ( − 0.0010)
0.5194 ( + 0.0002)
6.104 ( − 0.018)
7
0.3890 ( − 0.0051)
0.5212 ( + 0.0020)
6.134 ( + 0.013)
Base: default with energy_num_samples=4
Appendix
Table 6: Effect of energy_adaptive_weights defined in section A.3 , that replaces the uniform per-concept weights of the energy reward by a softmin over the normalized score-difference norms, on GenEval.
x0_hat_score_source
BLIP-VQA ↑
T2V ↑
HPSv3 ↑
Base: default
ϵtλ
0.3942
0.5192
6.122
ϵt §
0.3937 ( − 0.0004)
0.5180 ( − 0.0012)
6.109 ( − 0.013)
Appendix
Table 7: Effect of x0_hat_score_source , the score used to compute x^0 on GenEval.
num_ts_to_correct
BLIP-VQA ↑
T2V ↑
HPSv3 ↑
Base: default
5
0.3942
0.5192
6.122
2 §
0.3981 ( + 0.0040)
0.5087 ( − 0.0106)
6.140 ( + 0.019)
Base: default with eta=0.1 , nts_to_init_correct=0
5
0.3945
0.5110
6.137
10
0.3956 ( + 0.0011)
0.5120 ( + 0.0010)
6.126 ( − 0.011)
Appendix
Table 8: Effect of num_ts_to_correct , the number of corrected timesteps, on GenEval
num_latent_corrector_steps
BLIP-VQA ↑
T2V ↑
HPSv3 ↑
Base: default with guidance_scale=0.8 , eta=0.2 , energy_num_samples=4 ,
Table 9: Effect of num_latent_corrector_steps , the number of latent-corrector steps per corrected timestep (the first nts_to_init_correct timesteps use init_latent_corrector_steps instead), on GenEval.
We introduce the Noise-Tilted Reverse Kernel (NTRK), a reward-guided diffusion sampler that injects reward gradients through the noise term, leaving the pretrained reverse kernel unchanged and requiring only a single sample per step. Reward-guided sampling at inference time has greatly expanded the versatility of pretrained diffusion models. Yet existing methods face a trade-off. Gradient-based guidance shifts the reverse mean, steering generation but pushing intermediate states outside the region that the model was trained on and degrading quality. Search-based methods preserve quality but gain no gradient signal. No prior method achieves both. NTRK resolves this by keeping the reverse mean fixed and biasing the noise term toward high reward. This is enabled by a whitening operator, the central mechanism behind NTRK, which converts reward gradients into noise-compatible perturbations without losing their guiding signal. Across various reward alignment tasks, NTRK outperforms recent state-of-the-art baselines without losing sample quality. Remarkably, on aesthetic generation, NTRK surpasses the reward of the best baseline at 500 NFEs using only 25 NFEs, a 20 times reduction in compute.
Jisung Hwang, Yunhong Min, Jaihoon Kim +2
1KAIST · †This work was done in part while the author was a visiting researcher at The University of Tokyo. · 2The University of Tokyo
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary-condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify-then-Diffuse (RTD), a training-free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft-Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout-agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state-of-the-art compositional fidelity and robust gains. On the AE-Bench object pair subset, RTD improves BLIP-VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3× faster. Code will be released at https://github.com/Z-yiwei/rectify-then-diffuse
Ning Zhu, An Chen, Mengfei Zhao +4
Glasgow College, University of Electronic Science and Technology of China · School of Mathematical Sciences, University of Electronic Science and Technology of China
Compositional generalization requires models to produce novel configurations from familiar parts. In diffusion models, prior compositional generation methods typically assume that the relevant concepts or conditioning signals are already available. We instead ask whether a pretrained diffusion model can discover query-specific concepts from the time-indexed scores it learns for the noisy marginals pt(xt) and compose them at test time. Given a single out-of-distribution query, our method performs gradient ascent on sθ(xt,t)≈∇xtlogpt(xt) at multiple noising timesteps to recover local density modes, maps these modes into clean-space Gaussians, greedily selects relevant prototypes with a submodular likelihood objective, and combines them into a product-of-experts (PoE) teacher model with an analytic score. This teacher model can be sampled directly through classifier-free guidance or used to generate a sample pool for training a new class embedding and low-rank adapter. On held-out composition benchmarks built from ColorMNIST and CelebA, both the analytic PoE sampler and the low-rank adapted model outperform query-only and nearest trained-class baselines. These results suggest that the time-indexed score geometry of the diffusion model contains reusable density-mode concepts that support test-time compositional generation without a predefined concept library.
Zekun Wang, Anant Gupta, Tianyi Zhu +1
Georgia Institute of Technology · University of Virginia