From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
Authors: Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas
Organizations: ILIA, Université de Mons Mons, Belgium · Université Polytechnique Hauts-de-France Valenciennes, France · DeepILIA, ILIA, Université de Mons Mons, Belgium · ISIA Lab, Université de Mons Mons, Belgium
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom Δ-metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of −50% relative to real data, while CIA-only training shows a milder −16.5% degradation. Hybrid compositions significantly improve performance, with the 90% real + 10% Unity configuration achieving the best overall mAP@0.5 of 62.68% (+7.64% over baseline), and the 90% real + 10% CIA configuration maximizing precision at 74.45%. Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.
Figures & tables
Figure 1 . Overview of the proposed hybrid data pipeline. CIA generates photorealistic augmentations from real images using Diffusion, while Unity3D renders synthetic scenes with automatic annotations. The fused dataset Dmix balances realism and diversity, reducing the sim-to-real gap, and improving model robustness.
Figure 2 . Representative samples from the three dataset sources used in this study. Top: real construction-site images from MOCS. Middle: CIA-generated variants using diffusion-based controllable augmentation. Bottom: Unity-rendered synthetic scenes produced through domain randomization.
Group
Real
CIA
Unity
Precision
Recall
mAP@0.5
mAP@0.5:0.95
Fitness
Baselines
Real only
100
0
0
0.7266
0.4828
0.5504
0.3873
0.4036
CIA only
0
100
0
0.5923
0.3566
0.3851
0.2551
0.2681
Unity only
0
0
100
0.3004
0.0558
0.0447
0.0260
0.0279
Real–CIA Mixes
90
10
0
0.7445
0.5063
0.5792
0.4076
0.4248
Table 1 . Comprehensive performance comparison across all dataset composition experiments. Each configuration reports Precision, Recall, mAP@0.5, mAP@0.5:0.95, and Fitness, evaluated on a held-out real test set. Groups correspond to different mixing regimes between Real, Unity-simulated, and CIA-generated data. Green-highlighted row indicates the best overall result, while Blue-highlighted rows indicate the best results within each group. Purple-bordered cells mark the overall best value per metric, across all configurations.
Figure 3 . Non-normalized confusion matrices with a shared color scale. Raw counts expose the dominant error mode (FN→background) and the reduction achieved by small Unity/CIA mixes.
Figure 4 . Comparison of Δsim\textrightarrowreal and Δgen\textrightarrowreal across increasing synthetic data ratios