In interactive scenarios, an autonomous driving system is required to generate ego actions under the influence of other agents' behaviors. Existing World Action Models (WAMs) typically model other agents as components of the world model rather than as decision-makers that fundamentally shape the action of the ego agent, which impairs their performance in dense interaction scenarios. We introduce Reciprocal World Action Models (ReWAM), a game-theoretic world action modeling framework that captures the reciprocal influence between the ego agent and other agents by representing them as conditional responders whose actions are mutually influenced. We instantiate this framework with a Level-k response hierarchy, where role-specific ego and other action DiTs exchange compact strategy tokens through cross-agent attention while remaining grounded in a shared representation of the future driving world. To learn the response policy of the ego agent from demonstrations, we formulate expert actions as samples from the best response distribution and jointly optimize the entire hierarchy using conditional flow matching. Our framework is evaluated on the NAVSIM dataset and achieves state-of-the-art performance compared to baselines. The improvement is particularly significant in interactive scenarios, validating that modeling reciprocal responses provides a more effective foundation for interaction-aware world action generation.
Figures & tables
Figure 1: Comparison of different autonomous driving paradigms and representative interactive driving scenarios. (a) Vision-Language-Action models incorporate language-aligned semantic reasoning into action generation. (b) World Action Models ground ego-action generation in predicted world representations. (c) The proposed Reciprocal World Action Model explicitly captures reciprocal behavioral influence through recursive strategic reasoning. (d) Representative interactive scenarios in which the behaviors of the ego agent and other agents are mutually influenced.
Figure 2: The proposed ReWAM framework. Visual observations and language instructions are processed by the video branch to construct a shared latent representation of the future driving world. Conditioned on this representation and agent-specific state information, the Ego and Other Action DiTs synchronously generate their action hypotheses. At reasoning level k , each branch conditions on the level k−1 strategies of all remaining valid participants, forming a finite hierarchy of reciprocal responses.
Figure 3: Architecture of the proposed Level- k interaction block. Within each Action DiT block, temporal self-attention models the target agent’s action sequence, while cross-attention to the Video DiT features grounds action generation in the predicted driving world. Cross-agent attention further conditions the current response on the preceding-level action hypotheses of the remaining agents. The elaborated Action DiT architecture explicitly captures reciprocal strategic influence among agents in interactive scenarios.
Method
Ref
NC↑
DAC↑
TTC↑
Comf.↑
EP↑
PDMS↑
Traditional End-to-End Methods
VADv2- V8192 ( Jiang et al., 2024 )
arXiv’24
97.2
89.1
91.6
100
76.0
80.9
UniAD ( Hu et al., 2023 )
CVPR’23
97.8
91.9
92.9
100
78.8
83.4
TransFuser ( Chitta et al., 2023 )
TPAMI’23
97.7
92.8
92.8
100
79.2
84.0
PARA-Drive ( Weng et al., 2024 )
CVPR’24
97.9
92.4
93.0
99.8
79.3
84.0
DiffusionDrive ( Liao et al., 2025 )
CVPR’25
98.2
96.2
94.7
100
82.2
88.1
Table 1: Comparison with state-of-the-art methods on the NAVSIM benchmark. Higher values indicate better performance.
Interaction Level
NC
DAC
TTC
Comfort
EP
DDC
PDMS
Level-1
99.06
97.39
96.25
99.97
83.28
98.12
90.03
Level-2
99.25
97.68
97.13
99.97
83.44
98.17
90.48
Level-3
99.26
97.74
97.13
99.98
83.54
98.15
90.55
Level-4
99.25
97.81
97.18
99.96
83.60
98.14
90.64
Level-5
99.29
97.77
97.18
99.97
83.64
98.16
90.64
Table 2: The performance of the proposed method across different reasoning levels.
Interaction Intensity
Tokens
Share (%)
ReWAM
DriveLaW
DiffusionDrive
None
6621
54.52
89.92
88.40 (+1.52)
87.06 (+2.86)
Weak
1841
15.16
90.89
89.76 (+1.13)
89.20 (+1.69)
Medium
1841
15.16
92.33
91.73 (+0.60)
89.56 (+2.77)
Strong
1288
10.61
92.43
90.50 (+1.93)
90.17 (+2.26)
Critical
553
4.55
88.72
84.13 (+4.59)
86.18 (+2.54)
Table 3: PDMS scores across interaction-intensity groups on NAVTEST. Share denotes the percentage of the 12,144 evaluated tokens, and parentheses report ReWAM’s absolute improvement over each baseline.
Method
NC
DAC
TTC
Comfort
EP
DDC
PDMS
ReWAM
96.93
99.10
85.90
100
89.59
98.19
88.72
DriveLaW
94.21
98.55
79.20
100
87.06
98.01
84.13
DiffusionDrive
95.39
97.47
84.27
100
87.07
97.38
86.18
Table 4: Performance of the proposed method and baselines on critical-interaction scenarios.
Figure 4: Qualitative comparison across three representative interactive driving scenarios. Each panel presents the front-camera action projection, bird’s-eye-view trajectory visualization, and longitudinal speed profile.