Video generation for autonomous driving cannot follow the web-scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpus. We present a systematic scaling-law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters, trained at different exposures on up to 5,500 hours of driving. Validation loss follows consistent power laws in both model size and training exposure, answering the questions that shape a training budget: whether compute is better spent on longer training or on a larger model, and whether more data is needed. Loss improves much faster with training exposure than with model size, making longer training the most effective way to improve a fixed model under limited compute. However, larger models continue to achieve lower asymptotic loss, so compute-optimal scaling still favors increasing model size when sufficient compute and data are available. Guided by these laws, we train a 9B-parameter model, to our knowledge the largest video diffusion model trained from scratch on driving data: it sets a new open-source state of the art for driving video generation, as measured on nuScenes. Our code and pretrained models are available at https://github.com/valeoai/VATIX. NATIX is separately releasing the underlying driving data in stages.
Figures & tables
Model
Pico
Nano
Micro
Tiny
Small
Medium
Base
Big
Large
XLarge
1B
9B*
Params (M)
1.6
3.8
9.2
19.7
33.8
63.3
134.9
295.5
447.3
643.1
1142.9
9096.4
TFLOPs
0.02
0.05
0.14
0.30
0.61
1.10
2.47
5.61
8.79
12.99
23.09
143.84
Table 1 : Model size and training compute. * denotes the target model size. Training TFLOPs are estimated with thop as three times the FLOPs of a forward pass, approximating one training step (one forward and two backward-pass equivalents).
Figure 1 : Model architecture. Video latents from the Wan encoder are noised and processed by N transformer blocks, each containing spatial attention, temporal attention, and an MLP, modulated by AdaLN. The timestep t and ego-trajectory waypoints are embedded with sinusoidal features, and each is passed through an MLP. They are then summed and chunked into scale and shift. The output is the velocity (v-prediction).
Figure 2 : Model scaling analysis. Validation loss consistently decreases with increasing model size, and the fitted scaling law extrapolates this trend to larger models.
Figure 3 : Training dynamics per model size. Solid curves show measured validation loss during training; dashed curves show the corresponding per-model power law fit, extrapolated across the full exposure range shown; diamonds mark the best measured loss for each model. The dynamics are consistent across scales and well captured by per-model fits, while larger models converge to lower loss.
Figure 4 : Data restriction ablation. Validation loss at fixed training exposure (color) as the pool of unique footage available to the Base model is restricted from the full 5,500-hour corpus down to 5.5 hours. Loss degrades sharply only once repetition exceeds about 2,000 epochs (5.5h restriction); below that, additional unique footage brings diminishing returns for the Base model at the exposures tested here.
Figure 5 : Compute scaling analysis. Lower envelope of validation loss versus compute (a), and iso-FLOP optimal (N,D) configurations for each budget (b).
Figure 6 : Training-scaling fit for the 9B model. Validation loss during training and the corresponding power law fit ( L(D)=0.074799+0.00381D−0.7309 ), which suggests that additional training could still be beneficial to narrow the gap to the asymptotic loss L0 .
Model
FID Inception↓
FID DINO↓
FVD I3D↓
FVD VideoMAE↓
ADE ↓
Tiny
26.87
280.19
175.99
198.40
–
Base
10.66
132.51
68.43
116.94
–
Large
8.67
117.76
48.09
99.43
–
1B
5.61
87.05
33.94
89.49
–
9B
4.91
60.81
37.16
75.86
–
1B*
4.41
63.93
32.86
49.28
3.92 -0.98
Table 2 : Scaling Model size on 2.5 s video generation quality on NATIX , conditioned on a single frame. Lower values indicate better performance; best results in each column are bold. Fréchet distances are computed with different backbones. Models marked with * use trajectory-conditioning post-training. Green subscripts indicate the ADE boost over the corresponding unconditioned model. Larger models consistently improve generation quality, consistent with lower pre-training validation loss.
Model
FID Inception↓
FID DINO↓
FVD I3D↓
FVD VideoMAE↓
DriveDreamer-2 [ 40 ]
25.0
–
105.1
–
Drive-WM [ 35 ]
15.8
–
122.7
–
GenAD [ 37 ]
15.4
–
184.0
–
Vista [ 7 ]
6.9
–
89.4
–
GEM [ 9 ]
10.5
–
158.5
–
Driving World [ 15 ]
7.4
–
90.9
–
Table 3 : Comparison on the nuScenes benchmark. Quantitative evaluation of 2.5 s video generation conditioned on a single input frame (lower is better; “–” indicates the metric is not reported). Results are evaluated on two test splits matching the protocols of Vista (5,369 sequences) and Epona (1,690 sequences). Our models achieve state-of-the-art performance, demonstrating the benefit of scaling both model size and training data, followed by domain-specific fine-tuning on nuScenes.
Figure 7 : Visual Quality on 5-second generations. Small models (<20M parameters) capture the overall scene layout but quickly lose object consistency across frames. The base model (135M) better understands the scene, e.g., it correctly predicts that the car is turning left, but struggles to maintain coherent dynamics over longer horizons. Larger models (1.1B parameters) further improve visual quality and scene structure, yet still exhibit temporal inconsistencies and break down after approximately 2.5 seconds. Only our largest 9B models accurately capture the scene dynamics while maintaining high visual fidelity and temporal coherence throughout the entire rollout.