cs.CVAug 18, 2025

CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

Authors: Anindya Mondal, Sauradip Nag, Ayan Banerjee, Josep Llados, Xiatian Zhu, Anjan Dutta

Organizations: University of Surrey · Simon Fraser University · Universitat Autònoma de Barcelona

Abstract

Diffusion models excel at photorealistic synthesis but struggle with object count fidelity, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves structured instance and count control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism. Project page is at https://mondalanindya.github.io/CountLoop/.

Figures & tables

Explore similar work

CardsList
  1. Learning to Generate Multiple Objects from Dense and Occluded Layouts

    Jul 3, 2026Bach-Hoang Ngo, Si-Tri Ngo, Hieu Le +1Diverse Occlusion-And-Revelation ScenariosScene Reasoning

  2. ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation

    May 27, 2025Sanghyun Jo, Wooyeol Lee, Ziseok Lee +3Text-To-ImageDiffusion Models

  3. InstanceControl: Controllable Complex Image Generation without Instance Labeling

    Jun 30, 2026Xiaoyu Liu, Huan Wang, Fan Li +4Multi-Reference Image GenerationImage Generation