While Diffusion Models excel in text-to-image synthesis, they frequently suffer from catastrophic concept omission when generating complex multi-instance scenes. Existing training-free methods attempt to resolve this by rescaling attention maps, which merely exacerbates unstructured noise without establishing coherent semantic representations. To address this, we propose Delta-K, a backbone-agnostic, plug-and-play inference framework that resolves omission by operating directly in the shared cross-attention Key space. Utilizing a lightweight Vision-Language Model (VLM) preview, we isolate a differential key (ΔK) capturing the pure semantic signature of missing concepts, and proactively inject it during the early semantic planning phase. Governed by a dynamically optimized scheduling mechanism, Delta-K grounds diffuse noise into stable structural anchors while naturally preserving existing concepts via the inherent orthogonality of ΔK. Extensive experiments validate its universal applicability, demonstrating that Delta-K significantly improves compositional alignment across both modern DiT and foundational U-Net architectures without requiring spatial masks, auxiliary training, or structural modifications.
Figures & tables
Figure 1 : Spatiotemporal dynamics of attention in SD3.5. (a) Missing concepts suffer from chronic intensity suppression but follow valid temporal trends. (b) The high early AUC identifies a semantic planning phase for intervention before image structure solidifies. (c) High instability (CV) characterizes missing tokens as unstable noise.
Figure 2 : Overview of Delta-K. A VLM first separates present and missing concepts from a baseline generation. By contrasting the original and masked prompts, we obtain a differential key vector ΔK , which is dynamically injected into cross-attention keys during sampling to reinforce missing concepts while preserving existing content.
Figure 3 : Visualization. By using SD-2.1, Nano banana and DALL-E 3 Betker et al. [2023] as baseline methods for comparison, we observe that our approach achieves significant improvements in addressing the instance missing problem.
Figure 4 : Case study visualization. Left: Delta-K recovers omitted instances across SDXL and SD3.5; Middle: Cross-attention heatmaps for the SDXL example; Right: Evolution of attention weights (upper) and trajectories of CV and αt (lower).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1 : Attention entropy over denoising steps.
Figure B.1 : Spatiotemporal analysis of cross-attention dynamics in SDXL. (a) Intensity Divergence: Mean attention scores for “Present” vs. “Missing” concepts. (b) Early Detectability: AUC score for detecting omission; the curve remains relatively stable across steps with a peak AUC of 0.64. (c) Signal Stability: Coefficient of Variation (CV) showing that missing concepts correspond to unstable, high-variance attention patterns.
Figure G.1 : More examples of Delta-K compared with baseline and closed-source models.
Recent open-weight text-to-image (T2I) diffusion models still struggle with multi-instance prompts, often omitting or merging instances and mixing semantics among similar objects. We trace these failures to early denoising steps, before instance boundaries are reliably stabilized. Existing training-free guidance is largely driven by cross-attention or other token-conditioned semantic signals. Such guidance can separate concepts at the token level, but largely assumes that distinct instance regions have already emerged. In early denoising steps, it cannot reliably carve out these regions, so count failures and semantic mixing persist. By contrast, self-attention exposes class-agnostic instance layouts during early denoising. To exploit this asymmetry, we propose ISAC (Instance-to-Semantic Attention Control), a training-free, model-agnostic objective that first stabilizes self-attention layouts and then binds cross-attention semantics within them, without fine-tuning or external vision models. Across T2I-CompBench, HRS-Bench, and our newly curated IntraCompBench, ISAC consistently outperforms prior training-free methods. Furthermore, ISAC enhances layout-to-image controllers by refining coarse, overlapping bounding boxes into dense instance masks. Code and IntraCompBench are available at https://shjo-april.github.io/ISAC.
Sanghyun Jo, Wooyeol Lee, Ziseok Lee +3
OGQ, Seoul, Korea · Seoul National University, Seoul, Korea
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary-condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify-then-Diffuse (RTD), a training-free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft-Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout-agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state-of-the-art compositional fidelity and robust gains. On the AE-Bench object pair subset, RTD improves BLIP-VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3× faster. Code will be released at https://github.com/Z-yiwei/rectify-then-diffuse
Ning Zhu, An Chen, Mengfei Zhao +4
Glasgow College, University of Electronic Science and Technology of China · School of Mathematical Sciences, University of Electronic Science and Technology of China
Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omission, where specified objects or attributes fail to emerge in the generated image. By performing linear probing on text tokens, we demonstrate that text embeddings can distinguish a characteristic `omission signal' representing the absence of target concepts. Leveraging this insight, we propose Omission Signal Intervention (OSI), which amplifies the omission signal to actively catalyze the generation of missing concepts. Comprehensive experiments on FLUX.1-Dev and SD3.5-Medium demonstrate that OSI significantly alleviates concept omission even in extreme scenarios.
Kanghyun Baek, Jaihyun Lew, Chaehun Shin +2
Interdisciplinary Program in Artificial Intelligence, Seoul National University, Seoul, South Korea · Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea · Department of Computer Science & Engineering, Korea University, Seoul, South Korea +1