Driving world models provide a promising route toward scalable counterfactual data generation and interactive simulation beyond recorded driving logs. Realizing this potential requires a system that can generalize across diverse scenes, respond faithfully to prescribed controls, generate coherent multi-sensor observations, and operate efficiently under repeated inference. We present \textbf{HelloWorld}, a 2B driving world model system designed around these requirements. HelloWorld progressively specializes broad visual and motion priors from heterogeneous video data into controllable driving generation using ego pose, HD maps, and 3D boxes. A block-causal generation interface, together with adaptation to self-generated context, aligns the model with sequential simulation. The system further supports synchronized seven-camera RGB generation and conditional LiDAR synthesis, and is distilled toward few-step inference for efficient deployment. Experiments evaluate visual quality, control fidelity, cross-view consistency, robustness under repeated generation, inference efficiency, and LiDAR synthesis. Together, HelloWorld provides a unified framework for scalable driving data generation and interactive simulation.
Figures & tables
Figure 1 : Overview of HelloWorld. Top: A block-causal RGB generator is progressively specialized from pose-conditioned single-view modeling to synchronized seven-view synthesis and geometry-aware scene conditioning with ego pose, HD maps, and 3D boxes. The causal observation interface is retained across all stages. Bottom: A 20-step causal teacher is distilled into a four-step student, a decoupled super-resolution module enhances spatial fidelity, and an RGB-conditioned LiDAR branch synthesizes range and return intensity from synchronized seven-view observations.
Pose form
Multi- view
Structured control
Causal AR
Few- step
MagicDrive-V2 [ 13 ]
layout
✓
✓
✗
✗
GAIA-2 [ 17 ]
action
✓
✓
✓
✗
Cosmos-Transfer [ 15 ]
layout
△
✓
△
△
Cosmos WFM [ 2 ]
action
✗
△
✓
✓
LingBot-World [ 5 ]
action
✗
✗
✓
✓
HY-World 1.5 [ 4 ]
action
✗
✗
✓
✓
Table 1 : Capability comparison with representative world models. “Pose form” distinguishes independent metric 6-DoF ego pose, low-dimensional action, and ego motion implicit in layout conditioning. ✓ = supported, △ = partial/preliminary, ✗ = not reported. Entries reflect public reports as of Aug. 2026.
Figure 2 : Data construction under nested supervision. Pose-supervised video provides the broad pool for learning causal visual dynamics. Synchronized camera observations add multi-view geometry, and the structured-control subset supplies rendered maps and object annotations. Source-specific preparation and quality validation connect these supervision levels to the three training stages. The flows illustrate processing and supervision relationships rather than measured data proportions.
Source
Windows (w29)
Share
Views
Metric
HDMap/BBox
Fleet (7-camera)
1,283,914
35.1%
7
✓
subset
SpatialVID
991,073
27.1%
1
✗
✗
Sekai (real-walking)
746,509
20.4%
1
✗
✗
DL3DV
211,432
5.8%
1
✗
✗
ScanNet
185,548
5.1%
1
✗
✗
RealEstate10K
127,536
3.5%
1
✗
✗
Table 2 : Stage-1 training asset inventory (shares are pre-flatten window proportions, not effective sampler exposure): 29-frame training windows per source (754,629 sequences, 3,662,212 windows in total). Fleet windows are counted after caption-QA filtering of the 1,510,454 pose-balanced windows of Section 3.4 . With 7-View Flatten each fleet window yields 7 single-view samples.
Figure 3 : Overview of the HelloWorld generation framework. A causal teacher is distilled into a few-step student for seven-view RGB generation. A separate super-resolution module refines the generated views before conditional LiDAR synthesis. (a) The causal multi-view backbone incorporates ego-pose geometry through a ray encoder and AdaLN, scene layout through residuals from a control tower, and text through cross-attention, while cross-view attention exchanges information between neighboring cameras. (b) The RGB-to-LiDAR backbone conditions a LiDAR diffusion on encoded multi-view RGB latents while denoising LiDAR latents, which are decoded into range and intensity.
Method
Rot ↓
Trans ↓
Trans ∗ ↓
HY-WorldPlay
8.6305
1.0669
–
Lingbot World 2.0
3.2310
0.2853
–
HelloWorld
2.0620
0.2570
0.6830
Table 3 : Pose comparison against ground truth on the 29-frame single view benchmark. Rot and Trans use Sim(3) alignment. Trans ∗ denotes metric-scale translation error in meters, reported only for HelloWorld on real-world on-road vehicle cases. Lower is better for all metrics.
Method
Rot ↓
Trans ↓
Trans ∗ ↓
HY-WorldPlay
15.0835
5.4504
–
Lingbot World 2.0
5.4169
1.4988
–
HelloWorld
2.7056
1.3997
5.5857
Table 4 : Pose comparison against ground truth on the 100-frame single view benchmark. Rot and Trans use Sim(3) alignment. Trans ∗ denotes metric-scale translation error in meters, reported only for HelloWorld on real-world on-road vehicle cases. Lower is better for all metrics.
Method
SC
BC
MS
DD
AQ
IQ
I2V-S
I2V-B
Mean
HY-WorldPlay
0.9669
0.9543
0.9933
0.2800
0.4762
0.5774
0.9829
0.9847
0.7770
Lingbot World 2.0
0.9150
0.9424
0.9760
0.8900
0.4819
0.5933
0.9623
0.9676
0.8411
HelloWorld
0.9148
0.9364
0.9807
0.9200
0.4521
0.5392
0.9563
0.9632
0.8328
Table 5 : VBench comparison on the 29-frame single view benchmark. SC: subject consistency; BC: background consistency; MS: motion smoothness; DD: dynamic degree; AQ: aesthetic quality; IQ: imaging quality; I2V-S/B: image-to-video subject/background consistency. Mean is the unweighted arithmetic mean of the eight raw scores.
Method
SC
BC
MS
DD
AQ
IQ
I2V-S
I2V-B
Mean
HY-WorldPlay
0.9354
0.9325
0.9951
0.1400
0.4685
0.5584
0.9844
0.9855
0.7500
Lingbot World 2.0
0.8841
0.9270
0.9814
0.9900
0.4489
0.5816
0.9573
0.9624
0.8416
HelloWorld
0.8876
0.9288
0.9842
0.9200
0.4453
0.4976
0.9571
0.9627
0.8229
Table 6 : VBench comparison on the 100-frame single view benchmark.
Setting
Rot ( ∘ ) ↓
Trans ↓
Trans ∗↓
Cosmos Transfer 2.5 (w/ ff)
5.534
1.880
8.886
Stage 2 (w/ ff)
3.315
1.525
4.486
Stage 2 (w/o ff)
5.163
2.189
7.972
Stage 3 (w/ ff)
2.027
1.352
4.904
Stage 3 (w/o ff)
3.334
1.844
8.485
Table 7 : Pose comparison against ground truth on multiview benchmark.
Setting
CSE ↓
Δ CSE ↓
GT video
1.730
0.000
Cosmos Transfer 2.5 (w/ ff)
6.549
+4.819
Stage 2 (w/ ff)
6.192
+4.462
Stage 2 (w/o ff)
7.509
+5.779
Stage 3 (w/ ff)
4.180
+2.450
Stage 3 (w/o ff)
4.892
+3.162
Table 8 : Multiview CSE.
Setting
Subject
Background
Motion
Dynamic
Aesthetic
Imaging
Mean
Cosmos Transfer 2.5 (w/ ff)
0.8916
0.9376
0.9878
0.9555
0.4366
0.4239
0.7722
Stage 2 (w/ ff)
0.8999
0.9324
0.9843
0.9800
0.4374
0.4953
0.7882
Stage 2 (w/o ff)
0.8716
0.9199
0.9822
0.9700
0.4466
0.4844
0.7791
Stage 3 (w/ ff)
0.9089
0.9462
0.9876
0.9500
0.4138
0.4456
0.7754
Stage 3 (w/o ff)
0.9111
0.9472
0.9788
0.9700
0.4259
0.5074
0.7901
Table 9 : VBench comparison on the multiview benchmark.
Method
AR
MV
Video
DSF
FID ↓
FVD ↓
Short-horizon generation
DrivingGPT [ 37 ]
✗
✗
✓
✓
12.78
142.61
DrivingWorld [ 20 ]
✗
✗
✓
✓
7.40
90.90
Vista [ 18 ]
✗
✗
✓
✓
6.90
89.40
Epona [ 38 ]
✗
✗
✓
✓
7.50
82.80
BEVControl [ 8 ]
✗
✓
✗
✓
24.85
–
Table 10: Driving generation on nuScenes [ 36 ] at short and long horizons. AR: autoregressive generation; MV: multi-view; DSF: dense-supervision-free. Lower FID and FVD are better.
Figure 4 : Counterfactual pose control results. Given the same scene context, our model generates plausible future observations conditioned on counterfactual left-turn and right-turn ego-pose trajectories, demonstrating flexible and geometrically coherent control over ego motion.
Figure 5 : Unconventional lane-change control. We condition the model on two unconventional lane-change trajectories: a rightward lane change that crosses the roadside green belt into the non-motorized lane, and a leftward lane change that crosses the barrier into the opposite-direction lane. The generated results remain consistent with the prescribed pose conditions while preserving coherent scene structure.
Figure 6 : Comparison of scale control under different conditioning inputs. Our method supports both control video only and control video + pose. Compared with Cosmos transfer 2.5 and the control-video-only variant, adding pose leads to more accurate geometry, better scale consistency, and trajectories closer to the ground truth. Camera poses are estimated using Pi3X.
Figure 7 : Given the same underlying driving scenes, our model generates diverse weather conditions, including snowy, rainy, foggy, night-and-rainy, while preserving consistent scene structure and traffic semantics.
Figure 8 : Given the same underlying driving scenes, our model generates diverse weather conditions, including daylight, night, golden-hour, and blue-hour scenarios, while preserving consistent scene structure and traffic semantics.
Figure 9 : Special scene control results. Our model generates plausible long-tail traffic scenarios conditioned on special scene descriptions, demonstrating controllable synthesis of rare and safety-critical events while maintaining coherent scene structure and traffic context.
Setting
Rot ( ∘ ) ↓
Trans ↓
Trans ∗ ↓
Baseline (w/ ff)
2.027
1.352
4.904
Baseline (w/o ff)
3.334
1.844
8.485
dCM (w/ ff)
2.568
1.263
2.862
dCM (w/o ff)
4.619
2.907
5.311
DMD (w/ ff)
1.661
0.688
1.885
DMD (w/o ff)
2.311
1.034
3.193
Table 11: Multiview generation pose errors of few-step distilled models.
Setting
CSE ↓
Δ CSE ↓
GT video
1.730
0.000
Baseline (w/ ff)
4.180
+2.450
Baseline (w/o ff)
4.892
+3.162
dCM (w/ ff)
5.201
+3.471
dCM (w/o ff)
6.399
+4.669
DMD (w/ ff)
3.510
+1.780
Table 12: Multiview CSE of few-step distilled models.
Setting
Subject
Background
Motion
Dynamic
Aesthetic
Imaging
Mean
Baseline (w/ ff)
0.9089
0.9462
0.9876
0.9500
0.4138
0.4456
0.7754
Baseline (w/o ff)
0.9111
0.9472
0.9788
0.9700
0.4259
0.5074
0.7901
dCM (w/ ff)
0.8961
0.9409
0.9855
0.8800
0.4176
0.4281
0.7580
dCM (w/o ff)
0.9051
0.9469
0.9866
0.7600
0.4086
0.3824
0.7316
DMD (w/ ff)
0.9029
0.9430
0.9829
0.9900
0.4210
0.5039
0.7906
DMD (w/o ff)
0.9209
0.9488
0.9826
1.0000
0.4171
0.4935
0.7938
Table 13 : VBench metrics of few-step distilled models on multiview benchmark.
Phase
Teacher
DMD student
T/S
Text encode (s)
2.5
2.5
1.0×
VAE encode RGB (s)
6.7
7.1
0.95×
VAE encode control (s)
13.6
13.6
1.00×
VAE decode (s)
11.9
11.9
1.00×
DiT prefix (s)
1.1
1.7
0.68×
DiT ODE (s)
305.3
30.8
9.91×
Table 14 : Inference cost of the RF teacher and the DMD student.
Metric
Cosmos1
Ours
+ Temporal RRN
Range MAE (m) ↓
6.1724
2.2199
2.1629
Chamfer Distance (m) ↓
2.5709
1.0134
1.0426
Edge Pred → GT (m) ↓
1.5802
0.7976
0.7115
Ray-Align ↓
0.3776
0.4245
0.2757
Static-Align ↑
0.4503
0.6928
0.8463
Table 15: Common three-camera-region evaluation on 100 clips, frames 0–28. All five metrics use the same GT-derived output region; MAE uses pixels valid in all three predictions (82.7671% of regional valid GT). Static-Align uses shared reference-ICP transforms. Distances are in meters and alignment scores are fractions.
Metric
VAE
+ Single-frame RRN
+ Temporal RRN
Range MAE (m) ↓
0.3981
0.3839
0.3775
Chamfer Distance (m) ↓ [ 49 ]
0.3978
0.3815
0.3695
Edge Pred → GT (m) ↓ [ 50 ]
0.2997
0.2331
0.2106
Ray-Align ↓
0.2967
0.2610
0.2589
Static-Align ↑
0.8215
0.8510
0.8974
Warp Chamfer (m) ↓
0.4668
0.4422
0.3668
Table 16: LiDAR VAE reconstruction on 100 clips (2,900 frames). Range MAE uses common-valid returns; temporal scores use recorded poses. Distances are in meters and alignment scores are fractions.
Metric
Generator
+ Single-frame RRN
+ Temporal RRN
Range MAE (m) ↓
2.2797
2.2532
2.2194
Chamfer Distance (m) ↓
1.0738
1.1038
1.0920
Edge Pred → GT (m) ↓
0.8814
0.8339
0.7990
Ray-Align ↓
0.4366
0.3238
0.2822
Static-Align ↑
0.6741
0.7081
0.8310
Table 17: Full-panorama RGB-to-LiDAR generation on 100 clips (2,900 frames), without the three-camera-region restriction of Table 15 . All arms use the same saved generator outputs and reference data; rectification does not rerun the DiT. Distances are in meters and alignment scores are fractions.
Figure 10 : LiDAR VAE reconstruction across four recorded scenes: (a) overcast divided road, (b) daytime urban traffic, (c) illuminated tunnel, and (d) a wet roadway at night. For each scene, seven synchronized RGB views are followed by point clouds of GT, VAE reconstruction, VAE + single-frame RRN, and VAE + temporal RRN, from left to right. All examples use frame 14 and the same VAE. Point colors encode range with a shared 5–65 m color scale; this scale is not the evaluation range gate. RGB is shown only for scene context.
Figure 11 : RGB-to-LiDAR generation with temporal RRN on four recorded-road scenes. Each panel shows six displayed camera views, predicted range, auxiliary intensity, and a range-colored point cloud. Seven views condition generation; forward-far is omitted only from display. All four panels use the same temporal rectifier. Intensity is predicted by a separate head, not GT. These are generated and rectified point clouds, not VAE reconstructions.
Figure 12 : LiDAR synthesis conditioned on generated RGB: selected clear-morning and light-fog frames at t=2.0 s. Each panel pairs the displayed camera video mosaic with its synchronized rectified point cloud. These examples illustrate use beyond recorded RGB, not measured physical simulation of fog or a controlled weather ablation. Selected stills do not establish long-horizon consistency or temporal-RRN performance.
Figure 13 : Super-resolution enhancement. Our super-resolution module produces sharper boundaries and richer local details than bicubic upsampling, as highlighted by the zoomed-in lane-marking and vehicle regions.
Figure 14 : Special-scene editing results. Starting from recorded driving scenes, HelloWorld preserves the original road geometry, ego trajectory, and surrounding scene context while introducing rare traffic agents or events through edited semantic and object-level conditions. Each example shows the edited top-down BEV together with the corresponding synchronized multi-view observations. The edited scenarios include vehicle trajectory changes, object insertion, and safety-critical interactions involving pedestrians, motorbikes, and surrounding vehicles, demonstrating scene-consistent counterfactual editing under realistic driving contexts.
Figure 15 : Closed-loop simulation with Alpamayo 1.5 under trajectory adjustment. The left side shows the recorded road-test replay, while the right side shows the HelloWorld simulation driven by Alpamayo 1.5. At each step, Alpamayo 1.5 predicts an updated ego trajectory from HelloWorld-generated multi-view observations, and HelloWorld synthesizes the corresponding next observations. The rollout is executed with the few-step HelloWorld student model, enabling efficient repeated simulation. In this example, the simulated trajectory gradually deviates from the recorded path while remaining consistent with the surrounding road structure and scene context.
Figure 16 : Closed-loop simulation with Alpamayo 1.5 for obstacle avoidance. The left side shows the recorded road-test replay, while the right side shows the corresponding closed-loop simulation. Alpamayo 1.5 observes the multi-view frames generated by HelloWorld and predicts a lateral avoidance trajectory around a stopped vehicle. HelloWorld then generates the subsequent observations conditioned on the updated ego motion. The rollout uses the few-step HelloWorld student model and demonstrates interactive world-model simulation under action-model feedback.