Diffusion models excel at photorealistic synthesis but struggle with object count fidelity, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves structured instance and count control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism. Project page is at https://mondalanindya.github.io/CountLoop/.
Figures & tables
Figure 1 : Given prompts with explicit per-class counts, CountLoop produces images whose detected counts show substantially improved alignment with targets, even at high cardinalities ( e.g. , 140 oranges and 31 birds), where competing methods suffer from count saturation, semantic leakage, and grid-like layouts. Unlike prior approaches, CountLoop requires no retraining: a VLM-guided planning graph structures the layout, instance-driven attention masking prevents attribute leakage, and a Critic VLM iteratively refines the scene until the count and quality criteria are met. Count-faithful synthesis facilitates practical applications (right): (a) augmenting object-counting datasets ( Ranjan et al. (2021) ) with high-instance scenes, and (b) Compositional Image Generation in complex out-of-distribution scenarios for example it is rare to find peacock in the middle of traffic situation.
Figure 2 : Issues in High-instance image generation
Figure 3 : Given a text prompt, ⓐ The Design VLM parses the prompt to construct a planning graph, which is converted into a pixel-aligned layout ⓑ. ⓒ This layout guides an IP-Adapter-enhanced T2I backbone for image generation. ⓓ A Critic VLM evaluates the generated image’s count and aesthetics, providing structured feedback to update the planning graph. ⓔ This iterative loop continues until objectives are met.
Figure 4 : Cumulative latent composition, along with disentangled query feature extraction, mitigates attribute leakage.
Figure 5 : Successive layout refinement by Critic VLM. Layouts in the inset.
Figure 6 : CountLoop achieves significantly improved object count alignment and natural arrangements in dense scenes, while methods like LMD ( Lian et al. (2023) ), SLD ( Wu et al. (2024b) ), Counting Guidance ( Kang et al. (2025) ), and CountGen ( Binyamin et al. (2025) ) exhibit abnormal counts, spatial collapse, and grid artifacts. More visuals in the supplementary.
Family
Model
Single Category
Multi Category
COCO-Count
T2I-CompBench
CountLoop-S
CountLoop-M
MAE ↓
Ex. ↑
Sp. ↑
MAE ↓
Ex. ↑
Sp. ↑
MAE ↓
Ex. ↑
Sp. ↑
MAE ↓
Ex. ↑
Sp. ↑
SDXL ( Podell et al. (2024) )
2.37
25%
0.38
2.72
25%
0.75
29.96
0%
0.63
9.89
10%
0.55
FLUX ( Black Forest Labs (2024) )
1.40
55%
0.53
1.48
55%
0.78
17.47
4%
0.65
9.62
12%
0.58
SD 3.5 ( Stability AI (2024) )
1.10
55%
0.46
1.58
40%
0.76
21.81
0%
0.64
8.40
16%
0.56
SDXL-Turbo ( Sauer et al. (2024) )
2.50
25%
0.23
3.76
10%
0.53
51.14
0%
0.39
9.95
10%
0.37
Table 1 : Counting and spatial quality across benchmarks. We report counting error (MAE ↓ ), exact-count match (Ex. ↑ ), and spatial quality (Spatial ↑ ) for single-category and multi-category prompts. Agentic baselines are run with their native self-correction loops fully enabled. Best in bold, second-best underlined.
Figure 7 : Left: Counting error (MAE) rises with instance count. Right: Runtime curves echo the same ordering.
Table 2 : Analysis of CountLoop components, critic choices, and human evaluation. (a) Design-critic combinations on CountLoop-S ; the default is in bold and the best per designer is underlined . (b, c) Ablations of the design components ( PG : Planning Graph, CA : Cumulative Attn., IR : Iterative Refinement) and critic modules ( OVD : Open-vocab Detector, AS : Aesthetic Scorer). (d) User scores on a 0–5 scale (higher is better).
Backbone
MAE ↓
Spatial ↑
SD v1.5
8.05
0.88
SD 3.5
7.44
0.90
SDXL
7.59
0.93
Table 3 : Backbone swap.
Figure 8 : Per-regime MAE at representative instance counts on CountLoop-S . Categories are grouped by object size (Large / Medium / Small); bars show MAE at N∈{30,75,125,200} for all six methods. The Large < Medium < Small ordering is method-agnostic, reflecting canvas-density pressure and OWLv2 precision loss on small, densely-packed instances. CountLoop (red) achieves the lowest MAE in every regime at every N ; the gap over next-best 3DIS widens with N , reaching Δ MAE ≈7.6 at N=200 in the small-object regime (MAE 26.5 vs. 34.1, or 13% vs. 17% relative MAE). The dashed reference line marks CountLoop ’s overall MAE on CountLoop-S .
Critic OVD
Aesthetic
MAE ↓
vs Best
GroundingDINO
Q-Align (default)
7.59
+48%
OWLv2 †
Q-Align
8.57
+41%
GroundingDINO
ImageReward
9.21
+36%
Table 4 : Component swap on CountLoop-S . Evaluation detector is OWLv2 throughout. Values vs. Best Baseline report relative MAE reduction over 3DIS (MAE 14.55). † OWLv2 as both critic and evaluator gives an inherent advantage, yet MAE remains comparable.
Text-to-image diffusion models fail to generate correct object counts in dense scenes, where overlapping instances collapse into indistinguishable structures despite appearing visually plausible. We identify this as instance ownership collapse: tokens from overlapping objects interact freely through attention, while heavily occluded instances receive weak supervision due to their small visible areas. We address this through layout-aware attention biases that softly bias token interactions toward region-consistent grouping and suppress cross-instance leakage, paired with an amodal-balanced loss that amplifies gradients for occluded objects based on their occlusion level. To enable systematic evaluation, we introduce OverlapDepth-45K, a benchmark of densely overlapping scenes with amodal supervision. Our approach substantially improves count accuracy and prevents instance merging while preserving image quality. Project page: https://bachngoh.github.io/AIBL
Bach-Hoang Ngo, Si-Tri Ngo, Hieu Le +1
University of Science, Ho Chi Minh City, Vietnam · 2Vietnam National University, Ho Chi Minh City, Vietnam · 3UNC Charlotte, North Carolina, United States
Recent open-weight text-to-image (T2I) diffusion models still struggle with multi-instance prompts, often omitting or merging instances and mixing semantics among similar objects. We trace these failures to early denoising steps, before instance boundaries are reliably stabilized. Existing training-free guidance is largely driven by cross-attention or other token-conditioned semantic signals. Such guidance can separate concepts at the token level, but largely assumes that distinct instance regions have already emerged. In early denoising steps, it cannot reliably carve out these regions, so count failures and semantic mixing persist. By contrast, self-attention exposes class-agnostic instance layouts during early denoising. To exploit this asymmetry, we propose ISAC (Instance-to-Semantic Attention Control), a training-free, model-agnostic objective that first stabilizes self-attention layouts and then binds cross-attention semantics within them, without fine-tuning or external vision models. Across T2I-CompBench, HRS-Bench, and our newly curated IntraCompBench, ISAC consistently outperforms prior training-free methods. Furthermore, ISAC enhances layout-to-image controllers by refining coarse, overlapping bounding boxes into dense instance masks. Code and IntraCompBench are available at https://shjo-april.github.io/ISAC.
Sanghyun Jo, Wooyeol Lee, Ziseok Lee +3
OGQ, Seoul, Korea · Seoul National University, Seoul, Korea
Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with complex multi-instance scenes, frequently leading to attribute confusion among instances. While recent approaches attempt to mitigate this via manual instance labeling, such requirements are labor-intensive. In this paper, we propose InstanceControl, a novel multi-instance controllable generation method that eliminates the need for instance labeling. We identify the primary bottleneck in existing methods as the inability to accurately associate instance descriptions with their corresponding regions within visual conditions. To address this, we leverage the Vision-Language Model (VLM) to establish instance-level correspondences between text prompts and visual conditions. Specifically, the VLM automatically parses instance descriptions from the text prompts and simultaneously predicts instance masks based on the visual conditions. Furthermore, since the predicted masks may contain noise, we introduce an adaptive mask refinement strategy that dynamically refines these instance masks during the generation process. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods, achieving superior fidelity and precise instance-level control.
Xiaoyu Liu, Huan Wang, Fan Li +4
Harbin Institute of Technology, Harbin, China · HUAWEI Noah’s Ark Lab, Shenzhen, China