Organizations: Technical University of Munich, Germany; School of Engineering & Design, Department of Mobility Systems Engineering, Institute of Automotive Technology and Munich Institute of Robotics and Machine Intelligence (MIRMI) · MORAI Inc., Seoul, South Korea · AVP Division, Hyundai Motor Company, Gyeonggi-do, South Korea · Seoul National University, Seoul, South Korea
Camera-based 3D perception for autonomous driving relies heavily on large annotated datasets, and deploying such a system to a new target region typically requires data collection and annotation. Generative augmentation has been proposed to reduce this cost, but existing approaches face a fundamental trade-off: label-conditioned methods consume the very annotations they aim to replace, while simulator-conditioned methods offer free annotations but lack visual grounding to specific real environments. This work investigates the extent to which a digital-twin-driven Real2Sim2Real pipeline (DT-R2S2R) can substitute for target-region real data. By reconstructing recorded driving clips inside a georeferenced digital twin (DT-R2S), we condition a diffusion model on geometrically aligned simulator renderings, establishing a digital twin-grounded Sim2Real model (DT-S2R). As a result, DT-S2R synthesizes photorealistic driving images given low-cost yet georeferenced simulator data across both reconstructed and novel simulator scenes within digital-twin coverage. The efficacy of generated data is verified on diverse 3D detectors. DETR3D, especially, reports 93.18% of mAP obtained by a target-region real-data oracle, without employing target images for detector training. Furthermore, simple co-training with existing out-of-target real data outperforms the oracle. Thus, DT-R2S2R can substantially reduce the cost of manual on-site data collection and annotation in digital twin-available districts, providing a practical foundation for scaling 3D perception.
Figures & tables
Fig. 1: Overview of DT-R2S2R. Sim2Real model is trained on real driving images and reconstruction counterparts. Then the model uses newly synthesized simulator data to produce photorealistic images for training 3D detectors.
Fig. 2: Architecture of DT-R2S that locates a real driving clip inside the digital-twin simulator (1), replaces each object with the closest library asset (2), and recovers the LiDAR pose with respect to ego vehicle to improve real-sim image correspondence (3).
Fig. 3: Architecture of DT-S2R that drives one ControlNet branch per rendered modality (RGB, semantic, depth map) with cross-view blocks.
Fig. 4: Data generation using DT-S2R conditioned on both known and unknown driving routes from the simulator. (i) Geographical features and buildings are retained where the digital twin supplies them. (ii) Road-network detail survives DT-S2R. (iii) Cross-view consistency holds on a route absent from training.
Ablation
2D IoU ↑
3D Corner Err. ↓
SSIM ↑
LPIPS ↓
Baseline
57.5%
0.581
0.3427
0.6836
+ Appearance re-ranking
56.7%
0.632
0.3420
0.6849
+ Ground-plane refinement
59.4%
0.582
0.3430
0.6809
+ Clustering & Averaging
60.4%
0.565
0.3458
0.6804
TABLE II: Ablation study on DT-R2S. Best scores are shown in bold , and second-best scores are underlined .
Fig. 5: Visual comparison of translated images by injection mechanisms of rendered-RGB at Sim2Real step.
Inf. Input (val. set)
Mono-view
Multi-view
FCOS3D
DETR3D
PETR
mAP ↑
NDS ↑
mAP ↑
NDS ↑
mAP ↑
NDS ↑
DInTwinReal
52.2
57.4
53.4
57.6
60.4
61.0
DInTwinReal2Sim
24.8 (47.5%)
40.0 (69.6%)
30.8 (57.6%)
42.3 (73.4%)
31.1 (51.4%)
41.5 (68.0%)
DInTwinSim2Real
32.0 (61.3%)
45.6 (79.4%)
37.0 (69.2%)
47.5 (82.4%)
35.5 (58.7%)
46.8 (76.7%)
TABLE IV: Image fidelity quantified by FCOS3D trained on DInTwinReal . For validation clips of DInTwinReal , it takes each of the three rows as an input. The predictions by rows are scored against ground truth. Row colored blue denotes oracle.
Training set
mAP ↑
NDS ↑
DInTwinReal2Sim
6.86
27.31
DInTwinSim2Real
33.94
45.93
TABLE V: Detection performance of FCOS3D- w/o finetune trained for 12 epochs and evaluated on DAllReal .
Fig. 6: Preliminary attempt at long-tail mitigation utilizing DT-R2S2R. After reproducing a driving clip of DInTwinReal , a car showing normal behavior is edited to overtake and stop (red). The resulting images contain the augmented edge-case scenario, maintaining the original driving scenery (orange).