Diffusion models excel at photorealistic synthesis but struggle with object count fidelity, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves structured instance and count control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism. Project page is at https://mondalanindya.github.io/CountLoop/.
Figures & tables
Figure 1 : Given prompts with explicit per-class counts, CountLoop produces images whose detected counts show substantially improved alignment with targets, even at high cardinalities ( e.g. , 140 oranges and 31 birds), where competing methods suffer from count saturation, semantic leakage, and grid-like layouts. Unlike prior approaches, CountLoop requires no retraining: a VLM-guided planning graph structures the layout, instance-driven attention masking prevents attribute leakage, and a Critic VLM iteratively refines the scene until the count and quality criteria are met. Count-faithful synthesis facilitates practical applications (right): (a) augmenting object-counting datasets ( Ranjan et al. (2021) ) with high-instance scenes, and (b) Compositional Image Generation in complex out-of-distribution scenarios for example it is rare to find peacock in the middle of traffic situation.
Figure 2 : Issues in High-instance image generation
Figure 3 : Given a text prompt, ⓐ The Design VLM parses the prompt to construct a planning graph, which is converted into a pixel-aligned layout ⓑ. ⓒ This layout guides an IP-Adapter-enhanced T2I backbone for image generation. ⓓ A Critic VLM evaluates the generated image’s count and aesthetics, providing structured feedback to update the planning graph. ⓔ This iterative loop continues until objectives are met.
Figure 4 : Cumulative latent composition, along with disentangled query feature extraction, mitigates attribute leakage.
Figure 5 : Successive layout refinement by Critic VLM. Layouts in the inset.
Figure 6 : CountLoop achieves significantly improved object count alignment and natural arrangements in dense scenes, while methods like LMD ( Lian et al. (2023) ), SLD ( Wu et al. (2024b) ), Counting Guidance ( Kang et al. (2025) ), and CountGen ( Binyamin et al. (2025) ) exhibit abnormal counts, spatial collapse, and grid artifacts. More visuals in the supplementary.
Family
Model
Single Category
Multi Category
COCO-Count
T2I-CompBench
CountLoop-S
CountLoop-M
MAE ↓
Ex. ↑
Sp. ↑
MAE ↓
Ex. ↑
Sp. ↑
MAE ↓
Ex. ↑
Sp. ↑
MAE ↓
Ex. ↑
Sp. ↑
SDXL ( Podell et al. (2024) )
2.37
25%
0.38
2.72
25%
0.75
29.96
0%
0.63
9.89
10%
0.55
FLUX ( Black Forest Labs (2024) )
1.40
55%
0.53
1.48
55%
0.78
17.47
4%
0.65
9.62
12%
0.58
SD 3.5 ( Stability AI (2024) )
1.10
55%
0.46
1.58
40%
0.76
21.81
0%
0.64
8.40
16%
0.56
SDXL-Turbo ( Sauer et al. (2024) )
2.50
25%
0.23
3.76
10%
0.53
51.14
0%
0.39
9.95
10%
0.37
Table 1 : Counting and spatial quality across benchmarks. We report counting error (MAE ↓ ), exact-count match (Ex. ↑ ), and spatial quality (Spatial ↑ ) for single-category and multi-category prompts. Agentic baselines are run with their native self-correction loops fully enabled. Best in bold, second-best underlined.
Figure 7 : Left: Counting error (MAE) rises with instance count. Right: Runtime curves echo the same ordering.
Table 2 : Analysis of CountLoop components, critic choices, and human evaluation. (a) Design-critic combinations on CountLoop-S ; the default is in bold and the best per designer is underlined . (b, c) Ablations of the design components ( PG : Planning Graph, CA : Cumulative Attn., IR : Iterative Refinement) and critic modules ( OVD : Open-vocab Detector, AS : Aesthetic Scorer). (d) User scores on a 0–5 scale (higher is better).
Backbone
MAE ↓
Spatial ↑
SD v1.5
8.05
0.88
SD 3.5
7.44
0.90
SDXL
7.59
0.93
Table 3 : Backbone swap.
Figure 8 : Per-regime MAE at representative instance counts on CountLoop-S . Categories are grouped by object size (Large / Medium / Small); bars show MAE at N∈{30,75,125,200} for all six methods. The Large < Medium < Small ordering is method-agnostic, reflecting canvas-density pressure and OWLv2 precision loss on small, densely-packed instances. CountLoop (red) achieves the lowest MAE in every regime at every N ; the gap over next-best 3DIS widens with N , reaching Δ MAE ≈7.6 at N=200 in the small-object regime (MAE 26.5 vs. 34.1, or 13% vs. 17% relative MAE). The dashed reference line marks CountLoop ’s overall MAE on CountLoop-S .
Critic OVD
Aesthetic
MAE ↓
vs Best
GroundingDINO
Q-Align (default)
7.59
+48%
OWLv2 †
Q-Align
8.57
+41%
GroundingDINO
ImageReward
9.21
+36%
Table 4 : Component swap on CountLoop-S . Evaluation detector is OWLv2 throughout. Values vs. Best Baseline report relative MAE reduction over 3DIS (MAE 14.55). † OWLv2 as both critic and evaluator gives an inherent advantage, yet MAE remains comparable.
University of Science, Ho Chi Minh City, Vietnam · 2Vietnam National University, Ho Chi Minh City, Vietnam · 3UNC Charlotte, North Carolina, United States