Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video generation framework that jointly generates observations of vehicles sharing the same dynamic scene with precise camera-trajectory control. CoDrive interleaves local self-attention, which models spatiotemporal dependencies among the views of each vehicle, with global self-attention, which enables information exchange and consistency modeling across vehicles. To explicitly encode their spatial relationships, all camera trajectories are represented in a shared world coordinate system and injected into the attention layers through projective relative positional encoding. We further adopt a progressive mixed-task training strategy that combines large-scale real-world single-agent data with synthetic cross-agent interaction data, allowing the model to benefit from real-world appearance distributions while learning cross-agent consistency from simulation. For systematic evaluation, we introduce CoDrive-Bench, a benchmark covering real and synthetic multi-vehicle scenarios and evaluating trajectory controllability, scene geometry consistency, and instance-level consistency. Experiments show that CoDrive improves trajectory controllability and cross-agent geometric and instance consistency while maintaining competitive visual quality.
Figure 2: Model architecture and training framework overview. Given text prompt, reference frames, and camera poses from multiple agents and views, CoDrive encodes each video and jointly denoises them with a pretrained video DiT. Interleaved local self-attention models intra-agent spatial-temporal dependencies across multiple views, while global self-attention enables information exchange and consistency modeling across vehicles. Camera poses are expressed in a shared world coordinate system and injected into both attention modules through projective relative positional encoding (PRoPE), providing explicit geometric guidance for cross-view and cross-agent alignment.
Figure 3: CoDrive-Bench data statistics. (a) Distribution of inter-agent distances in the dataset. (b) Distribution of co-visibility between the two agents. (c) Word cloud of text prompts.
Controllability
Scene Consistency
Instance Consistency
Method
FVD ↓
ADE ↓
DTW ↓
CD ↓
RE ↓
Hit ↑
LE ↓
HunyuanVideo-1.5
1082.4
11.96
573
2.08
0.61
0.322
10.92
Wan2.2-I2V
444.2
22.95
1189
1.82
0.64
0.302
11.16
MagicDrive-V2
2815.4
10.71
546
19.53
1.27
0.007
12.00
ShareVerse
228.0
15.25
399
4.21
1.11
0.159
9.40
Cosmos3
103.2
5.69
241
3.62
1.27
0.373
9.83
Table 1: Quantitative comparison of video generation methods. ↑ indicates that higher values are better, while ↓ indicates that lower values are better.
Controllability
Scene Consistency
Instance Consistency
IGLA
GCGI
FVD ↓
ADE ↓
DTW ↓
CD ↓
RE ↓
Hit ↑
LE ↓
✓
161.2
6.71
321
1.77
0.59
0.409
7.25
✓
238.1
13.40
688
1.75
0.55
0.342
8.28
✓
✓
108.8
3.05
116
1.64
0.52
0.482
5.81
Table 2: Component ablation. Component ablation on CoDrive-Bench. We isolate Interleaved Global-Local Self-Attention (IGLA) and Global Camera Geometry Injection (GCGI) respectively and verify their complementary effects.
Controllability
Scene Consistency
Instance Consistency
FVD ↓
ADE ↓
DTW ↓
CD ↓
RE ↓
Hit ↑
LE ↓
w/o Mixed Fine-tuning
282.9
6.86
343
1.68
0.47
0.560
7.69
w/ Mixed Fine-tuning
104.3
1.53
61
1.57
0.45
0.646
5.99
Table 3: Impact of mixed-task fine-tuning on performance in real-world scenarios.
Figure 4: Qualitative results on both real-world and synthetic data. Blue image borders indicate Agent A’s three views, while orange image borders indicate Agent B’s three views. Blue bounding boxes mark Agent A in Agent B’s views, whereas orange bounding boxes mark Agent B in Agent A’s views.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Interaction mode
Description
Intersection crossing
The two vehicles approach a common region from substantially different directions and interact near an intersection or road crossing.
Same-direction following
The two vehicles travel in approximately the same direction while maintaining a relatively short longitudinal distance.
Oncoming passing
The two vehicles approach one another with nearly opposite headings and pass each other on opposing lanes or nearby roads.
Turning interaction
At least one of the two vehicles undergoes a significant heading change while the vehicles remain spatially close.
Merging or cut-in
The vehicles have approximately aligned headings and become substantially closer over time, covering road merging, lane convergence, and cut-in-like interactions.
Parallel-lane driving
The two vehicles travel on approximately parallel roads or lanes without satisfying the stronger geometric conditions of the other interaction modes.
Appendix
Table 4: Vehicle interaction modes used in the synthetic data collection pipeline.
Synthetic
Real
count
share
count
share
Not resolvable by the camera rig
Target behind the observer
10,802
29.5%
13,255
36.2%
Projects to <0.3% of image
14,020
38.3%
14,606
39.9%
Geometrically visible
11,778
32.2%
8,739
23.9%
Excluded by the metric, as a share of the visible set
Appendix
Table 5: The scored observation set is constructed from pose-based criteria that apply identically to all methods, yielding an exactly paired comparison free of gating bias.
Figure 5: Sensitivity analysis of CoDrive-Bench consistency metrics under controlled perturbations . We progressively perturb the depth estimates and participating-agent position starting from the unperturbed reference setting. Increasing depth noise monotonically increases CD and RE, while increasing agent-position perturbation decreases Soft Hit Rate and increases LE. The annotations report the relative change between the largest perturbation and the unperturbed setting.
Metric
Evaluator Variants
Spearman ρ
CoDrive Rank
CD
UniDepth vs. VideoDepthAnything
1.00
#1
RE
UniDepth vs. VideoDepthAnything
0.94
#1
Hit
ByteTrack vs. BoT-SORT
1.00
#1
LE
ByteTrack vs. BoT-SORT
1.00
#1
Appendix
Table 6: Robustness of CoDrive-Bench metrics to evaluator choice . We compare model rankings obtained using alternative depth estimators for CD and RE, and alternative multi-object trackers for Hit and LE. Spearman’s ρ is computed across all evaluated methods. Rankings remain highly consistent across evaluator variants, and CoDrive retains the top rank in all cases.
Controllability
Scene Consistency
Instance Consistency
Method
FVD ↓
ADE ↓
DTW ↓
CD ↓
RE ↓
Hit ↑
LE ↓
HunyuanVideo-1.5
882.3
11.20
599
2.04
0.53
0.527
8.87
Wan2.2-I2V
850.5
27.13
1524
1.81
0.62
0.470
8.82
Cosmos3-Nano
106.4
2.60
117
3.40
1.33
0.632
6.71
MagicDrive-V2
2277.0
5.29
270
13.19
1.42
0.013
6.43
ShareVerse
416.6
19.65
544
3.76
1.01
0.290
7.29
Appendix
Table 7: Quantitative comparison of video generation methods on real-world data.
Controllability
Scene Consistency
Instance Consistency
Method
FVD ↓
ADE ↓
DTW ↓
CD ↓
RE ↓
Hit ↑
LE ↓
HunyuanVideo-1.5
1676.8
12.71
548
2.11
0.69
0.118
12.97
Wan2.2-I2V
283.0
18.77
855
1.82
0.66
0.134
13.49
Cosmos3-Nano
205.8
8.78
366
3.85
1.22
0.115
12.96
MagicDrive-V2
3468.9
16.1
822
25.86
1.13
0.001
19.57
ShareVerse
240.0
10.86
254
4.66
1.20
0.027
11.51
Appendix
Table 8: Quantitative comparison of video generation methods on synthetic data.
Figure 6: Additional qualitative results . Each row shows the relative trajectories of Agents A and B together with selected generated frames from both agents.