cs.CVMar 10, 2026

Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation

Authors: Zitong Wang, Zijun Shen, Haohao Xu, Zhengjie Luo, Weibin Wu

Organizations: School of Software Engineering, Sun Yat-sen University · Nanjing University · College of Management and Economics, Tianjin University

Abstract

While Diffusion Models excel in text-to-image synthesis, they frequently suffer from catastrophic concept omission when generating complex multi-instance scenes. Existing training-free methods attempt to resolve this by rescaling attention maps, which merely exacerbates unstructured noise without establishing coherent semantic representations. To address this, we propose Delta-K, a backbone-agnostic, plug-and-play inference framework that resolves omission by operating directly in the shared cross-attention Key space. Utilizing a lightweight Vision-Language Model (VLM) preview, we isolate a differential key (ΔKΔK) capturing the pure semantic signature of missing concepts, and proactively inject it during the early semantic planning phase. Governed by a dynamically optimized scheduling mechanism, Delta-K grounds diffuse noise into stable structural anchors while naturally preserving existing concepts via the inherent orthogonality of ΔKΔK. Extensive experiments validate its universal applicability, demonstrating that Delta-K significantly improves compositional alignment across both modern DiT and foundational U-Net architectures without requiring spatial masks, auxiliary training, or structural modifications.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation

    May 27, 2025Sanghyun Jo, Wooyeol Lee, Ziseok Lee +3Text-To-ImageDiffusion Models

  2. Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds

    Aug 4, 2026Ning Zhu, An Chen, Mengfei Zhao +4Text-To-Image Diffusion ModelsDenoising Trajectory

  3. Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers

    May 14, 2026Kanghyun Baek, Jaihyun Lew, Chaehun Shin +2Diffusion TransformersOmissions