In interactive scenarios, an autonomous driving system is required to generate ego actions under the influence of other agents' behaviors. Existing World Action Models (WAMs) typically model other agents as components of the world model rather than as decision-makers that fundamentally shape the action of the ego agent, which impairs their performance in dense interaction scenarios. We introduce Reciprocal World Action Models (ReWAM), a game-theoretic world action modeling framework that captures the reciprocal influence between the ego agent and other agents by representing them as conditional responders whose actions are mutually influenced. We instantiate this framework with a Level-k response hierarchy, where role-specific ego and other action DiTs exchange compact strategy tokens through cross-agent attention while remaining grounded in a shared representation of the future driving world. To learn the response policy of the ego agent from demonstrations, we formulate expert actions as samples from the best response distribution and jointly optimize the entire hierarchy using conditional flow matching. Our framework is evaluated on the NAVSIM dataset and achieves state-of-the-art performance compared to baselines. The improvement is particularly significant in interactive scenarios, validating that modeling reciprocal responses provides a more effective foundation for interaction-aware world action generation.
Figures & tables
Figure 1: Comparison of different autonomous driving paradigms and representative interactive driving scenarios. (a) Vision-Language-Action models incorporate language-aligned semantic reasoning into action generation. (b) World Action Models ground ego-action generation in predicted world representations. (c) The proposed Reciprocal World Action Model explicitly captures reciprocal behavioral influence through recursive strategic reasoning. (d) Representative interactive scenarios in which the behaviors of the ego agent and other agents are mutually influenced.
Figure 2: The proposed ReWAM framework. Visual observations and language instructions are processed by the video branch to construct a shared latent representation of the future driving world. Conditioned on this representation and agent-specific state information, the Ego and Other Action DiTs synchronously generate their action hypotheses. At reasoning level k , each branch conditions on the level k−1 strategies of all remaining valid participants, forming a finite hierarchy of reciprocal responses.
Figure 3: Architecture of the proposed Level- k interaction block. Within each Action DiT block, temporal self-attention models the target agent’s action sequence, while cross-attention to the Video DiT features grounds action generation in the predicted driving world. Cross-agent attention further conditions the current response on the preceding-level action hypotheses of the remaining agents. The elaborated Action DiT architecture explicitly captures reciprocal strategic influence among agents in interactive scenarios.
Method
Ref
NC↑
DAC↑
TTC↑
Comf.↑
EP↑
PDMS↑
Traditional End-to-End Methods
VADv2- V8192 ( Jiang et al., 2024 )
arXiv’24
97.2
89.1
91.6
100
76.0
80.9
UniAD ( Hu et al., 2023 )
CVPR’23
97.8
91.9
92.9
100
78.8
83.4
TransFuser ( Chitta et al., 2023 )
TPAMI’23
97.7
92.8
92.8
100
79.2
84.0
PARA-Drive ( Weng et al., 2024 )
CVPR’24
97.9
92.4
93.0
99.8
79.3
84.0
DiffusionDrive ( Liao et al., 2025 )
CVPR’25
98.2
96.2
94.7
100
82.2
88.1
Table 1: Comparison with state-of-the-art methods on the NAVSIM benchmark. Higher values indicate better performance.
Interaction Level
NC
DAC
TTC
Comfort
EP
DDC
PDMS
Level-1
99.06
97.39
96.25
99.97
83.28
98.12
90.03
Level-2
99.25
97.68
97.13
99.97
83.44
98.17
90.48
Level-3
99.26
97.74
97.13
99.98
83.54
98.15
90.55
Level-4
99.25
97.81
97.18
99.96
83.60
98.14
90.64
Level-5
99.29
97.77
97.18
99.97
83.64
98.16
90.64
Table 2: The performance of the proposed method across different reasoning levels.
Interaction Intensity
Tokens
Share (%)
ReWAM
DriveLaW
DiffusionDrive
None
6621
54.52
89.92
88.40 (+1.52)
87.06 (+2.86)
Weak
1841
15.16
90.89
89.76 (+1.13)
89.20 (+1.69)
Medium
1841
15.16
92.33
91.73 (+0.60)
89.56 (+2.77)
Strong
1288
10.61
92.43
90.50 (+1.93)
90.17 (+2.26)
Critical
553
4.55
88.72
84.13 (+4.59)
86.18 (+2.54)
Table 3: PDMS scores across interaction-intensity groups on NAVTEST. Share denotes the percentage of the 12,144 evaluated tokens, and parentheses report ReWAM’s absolute improvement over each baseline.
Method
NC
DAC
TTC
Comfort
EP
DDC
PDMS
ReWAM
96.93
99.10
85.90
100
89.59
98.19
88.72
DriveLaW
94.21
98.55
79.20
100
87.06
98.01
84.13
DiffusionDrive
95.39
97.47
84.27
100
87.07
97.38
86.18
Table 4: Performance of the proposed method and baselines on critical-interaction scenarios.
Figure 4: Qualitative comparison across three representative interactive driving scenarios. Each panel presents the front-camera action projection, bird’s-eye-view trajectory visualization, and longitudinal speed profile.
A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action Models (WAMs) largely miss this reciprocity, treating world prediction and action generation as either isolated parallel branches or rigid predict-then-plan pipelines. We formalize this perspective as World-Action Interactive Models (WAIMs), and instantiate it in autonomous driving with \textbf{DAWN} (\textbf{D}enoising \textbf{A}ctions and \textbf{W}orld i\textbf{N}teractive model), a simple yet strong latent generative baseline. DAWN operates in a compact semantic latent space and couples a \emph{World Predictor} with a \emph{World-Conditioned Action Denoiser}: the predicted world hypothesis conditions action denoising, while the denoised action hypothesis is fed back to update the world prediction, so that both are recursively refined during inference. Rather than eliminating test-time world evolution altogether or rolling out the full future in pixel space, DAWN performs a short explicit latent rollout that is sufficient to support long-horizon trajectory generation in complex interactive scenes. Experiments show that DAWN achieves strong planning performance and favorable safety-related results across multiple autonomous driving benchmarks. More broadly, our results suggest that interactive world-action generation is a principled path toward truly actionable world models.
Hongbo Lu, Liang Yao, Chenghao He +6
1COWARobot Co. Ltd · 2Shanghai Jiao Tong University · 3Hohai University
In autonomous driving, World-Action Models (WAMs) have improved end-to-end planning by transferring video dynamics priors to action prediction, but many still couple planning with future-video generation at inference, incurring substantial computational overhead. We present SimWAM, a simple yet effective WAM that leverages future-video prediction solely as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without future-frame generation at inference. This design supports multiple pretrained video backbones and independent action-expert scaling within a shared attention interface, while preserving the joint learning objective. Moreover, we apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Experiments show that SimWAM achieves 91.9 PDMS on NAVSIM with a favorable trade-off between accuracy and latency among world-model-based planners, while transferring zero-shot to nuScenes. It also achieves competitive planning accuracy on WOD-E2E and PhysicalAI-Autonomous-Vehicles. These results position SimWAM as a plain yet solid baseline for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.
Zongchuang Zhao, Xin Zhou, Tianyang Xu +6
Huazhong University of Science & Technology · Dongfeng Research & Development Institute
Modern driving action models are increasingly improved in a self-improvement loop, where a learned world simulator imagines future observations and the resulting data is fed back to refine the action model. However, the bottleneck of this loop lies in the simulators' inability to generate behaviorally plausible responses by surrounding agents, making generated data both unrealistic in interaction and imbalanced in distribution. We introduce BehaviorWorldGen, a framework that closes the loop between action models and world simulators through controllable behavior-aware structured world generation. Its core component is BehaviorFlow, a meta-action-conditioned traffic-flow model that injects interpretable behavior controls and jointly generates multi-agent rollouts. BehaviorFlow realizes the specified agent behaviors while allowing surrounding vehicles to respond to the ego and to one another. The resulting rollouts are rendered by a world simulator into realistic multi-view observations, which are paired with corrected interaction-aware trajectories for action-model refinement. Since BehaviorWorldGen uses structured trajectories as the interface between its modules, it is compatible with diverse action models and world simulators. Experiments on world generation, scene extrapolation, and policy refinement demonstrate consistent improvements, with the largest benefits concentrated on difficult interactive scenarios.