From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
Organizations: ILIA, Université de Mons Mons, Belgium · Université Polytechnique Hauts-de-France Valenciennes, France · DeepILIA, ILIA, Université de Mons Mons, Belgium · ISIA Lab, Université de Mons Mons, Belgium
Abstract
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom -metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of relative to real data, while CIA-only training shows a milder degradation. Hybrid compositions significantly improve performance, with the 90% real + 10% Unity configuration achieving the best overall mAP@0.5 of ( over baseline), and the 90% real + 10% CIA configuration maximizing precision at . Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.
Figures & tables
| Group | Real | CIA | Unity | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 | Fitness |
| Baselines | ||||||||
| Real only | 0.7266 | 0.4828 | 0.5504 | 0.3873 | 0.4036 | |||
| CIA only | ||||||||
| Unity only | ||||||||
| Real–CIA Mixes | ||||||||
| 0.7445 | 0.5063 | 0.5792 | 0.4076 | 0.4248 |