Organizations: Technical University of Munich, Germany; School of Engineering & Design, Department of Mobility Systems Engineering, Institute of Automotive Technology and Munich Institute of Robotics and Machine Intelligence (MIRMI) · MORAI Inc., Seoul, South Korea · AVP Division, Hyundai Motor Company, Gyeonggi-do, South Korea · Seoul National University, Seoul, South Korea
Camera-based 3D perception for autonomous driving relies heavily on large annotated datasets, and deploying such a system to a new target region typically requires data collection and annotation. Generative augmentation has been proposed to reduce this cost, but existing approaches face a fundamental trade-off: label-conditioned methods consume the very annotations they aim to replace, while simulator-conditioned methods offer free annotations but lack visual grounding to specific real environments. This work investigates the extent to which a digital-twin-driven Real2Sim2Real pipeline (DT-R2S2R) can substitute for target-region real data. By reconstructing recorded driving clips inside a georeferenced digital twin (DT-R2S), we condition a diffusion model on geometrically aligned simulator renderings, establishing a digital twin-grounded Sim2Real model (DT-S2R). As a result, DT-S2R synthesizes photorealistic driving images given low-cost yet georeferenced simulator data across both reconstructed and novel simulator scenes within digital-twin coverage. The efficacy of generated data is verified on diverse 3D detectors. DETR3D, especially, reports 93.18% of mAP obtained by a target-region real-data oracle, without employing target images for detector training. Furthermore, simple co-training with existing out-of-target real data outperforms the oracle. Thus, DT-R2S2R can substantially reduce the cost of manual on-site data collection and annotation in digital twin-available districts, providing a practical foundation for scaling 3D perception.
Figures & tables
Fig. 1: Overview of DT-R2S2R. Sim2Real model is trained on real driving images and reconstruction counterparts. Then the model uses newly synthesized simulator data to produce photorealistic images for training 3D detectors.
Fig. 2: Architecture of DT-R2S that locates a real driving clip inside the digital-twin simulator (1), replaces each object with the closest library asset (2), and recovers the LiDAR pose with respect to ego vehicle to improve real-sim image correspondence (3).
Fig. 3: Architecture of DT-S2R that drives one ControlNet branch per rendered modality (RGB, semantic, depth map) with cross-view blocks.
Fig. 4: Data generation using DT-S2R conditioned on both known and unknown driving routes from the simulator. (i) Geographical features and buildings are retained where the digital twin supplies them. (ii) Road-network detail survives DT-S2R. (iii) Cross-view consistency holds on a route absent from training.
Ablation
2D IoU ↑
3D Corner Err. ↓
SSIM ↑
LPIPS ↓
Baseline
57.5%
0.581
0.3427
0.6836
+ Appearance re-ranking
56.7%
0.632
0.3420
0.6849
+ Ground-plane refinement
59.4%
0.582
0.3430
0.6809
+ Clustering & Averaging
60.4%
0.565
0.3458
0.6804
TABLE II: Ablation study on DT-R2S. Best scores are shown in bold , and second-best scores are underlined .
Fig. 5: Visual comparison of translated images by injection mechanisms of rendered-RGB at Sim2Real step.
Inf. Input (val. set)
Mono-view
Multi-view
FCOS3D
DETR3D
PETR
mAP ↑
NDS ↑
mAP ↑
NDS ↑
mAP ↑
NDS ↑
DInTwinReal
52.2
57.4
53.4
57.6
60.4
61.0
DInTwinReal2Sim
24.8 (47.5%)
40.0 (69.6%)
30.8 (57.6%)
42.3 (73.4%)
31.1 (51.4%)
41.5 (68.0%)
DInTwinSim2Real
32.0 (61.3%)
45.6 (79.4%)
37.0 (69.2%)
47.5 (82.4%)
35.5 (58.7%)
46.8 (76.7%)
TABLE IV: Image fidelity quantified by FCOS3D trained on DInTwinReal . For validation clips of DInTwinReal , it takes each of the three rows as an input. The predictions by rows are scored against ground truth. Row colored blue denotes oracle.
Training set
mAP ↑
NDS ↑
DInTwinReal2Sim
6.86
27.31
DInTwinSim2Real
33.94
45.93
TABLE V: Detection performance of FCOS3D- w/o finetune trained for 12 epochs and evaluated on DAllReal .
Fig. 6: Preliminary attempt at long-tail mitigation utilizing DT-R2S2R. After reproducing a driving clip of DInTwinReal , a car showing normal behavior is edited to overtake and stop (red). The resulting images contain the augmented edge-case scenario, maintaining the original driving scenery (orange).
Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D vehicle generation methods are often trained on synthetic data with significant domain gaps from real-world distributions. The generated models often exhibit arbitrary poses and undefined scales, resulting in poor visual consistency when integrated into driving scenes. In this paper, we present Unposed-to-3D, a novel framework that learns to reconstruct 3D vehicles from real-world driving images using image-only supervision. Our approach consists of two stages. In the first stage, we train an image-to-3D reconstruction network using posed images with known camera parameters. In the second stage, we remove camera supervision and use a camera prediction head that directly estimates the camera parameters from unposed images. The predicted pose is then used for differentiable rendering to provide self-supervised photometric feedback, enabling the model to learn 3D geometry purely from unposed images. To ensure simulation readiness, we further introduce a scale-aware module to predict real-world size information, and a harmonization module that adapts the generated vehicles to the target driving scene with consistent lighting and appearance. Extensive experiments demonstrate that Unposed-to-3D effectively reconstructs realistic, pose-consistent, and harmonized 3D vehicle models from real-world images, providing a scalable path toward creating high-quality assets for driving scene simulation and digital twin environments.
Hongyuan Liu, Bochao Zou, Qiankun Liu +10
University of Science and Technology Beijing · Li Auto Inc.
Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated data, the domain gap between synthetic and real-world driving videos significantly limits its utility for downstream deployment. Existing video generation methods are not well-suited for this task, as they fail to simultaneously preserve scene structure, object dynamics, temporal consistency, and visual realism, all of which are critical for maintaining annotation validity in generated data. In this paper, we present DriveCtrl, a depth-conditioned controllable sim-to-real video generation framework for realistic driving video synthesis. Built upon a pretrained video foundation model, DriveCtrl introduces a structure-aware adapter that enables depth-guided generation while preserving the scene layout and motion patterns of the source simulation, producing temporally coherent driving videos that remain aligned with the original simulated sequences. We further introduce a scalable data generation pipeline that transforms simulator videos into realistic driving footage matching the visual style of a target real-world dataset. The pipeline supports three conditioning signals: structural depth, reference-dataset style, and text prompts, while preserving frame-level annotations for downstream perception tasks. To better assess this task, we propose a driving-domain-specific knowledge-informed evaluation metric called Driving Video Realism Score (DVRS) that assesses the realism of generated videos. Experiments demonstrate that DriveCtrl consistently outperforms the base model and competing alternatives in realism, temporal quality, and perception task performance, substantially narrowing the sim-to-real gap for driving video generation.
Haonan Zhao, Yiting Wang, Jingkun Chen +3
University of Warwick · 2Northwestern Polytechnical University · 3Queen Mary University of London
High-quality 3D assets for traffic participants are critical for multi-sensor simulation, which is essential for the safe end-to-end development of autonomy. Building assets from in-the-wild data is key for diversity and realism, but existing neural-rendering based reconstruction methods are slow and generate assets that render well only from viewpoints close to the original observations, limiting their usefulness in simulation. Recent diffusion-based generative models build complete and diverse assets, but perform poorly on in-the-wild driving scenes, where observed actors are captured under sparse and limited fields of view, and are partially occluded. In this work, we propose a 3D latent diffusion model that learns on in-the-wild LiDAR and camera data captured by a sensor platform and generates high-quality 3D assets with complete geometry and appearance. Key to our method is a "reconstruct-then-generate" approach that first leverages occlusion-aware neural rendering trained over multiple scenes to build a high-quality latent space for objects, and then trains a diffusion model that operates on the latent space. We show our method outperforms existing reconstruction and generation based methods, unlocking diverse and scalable content creation for simulation.