World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving
Authors: Jieyuan Pei, Meiyi Lu, Sining Ang, Yubo Zhao, Zhangyi Hu, Mingwei Xu, Haokai Ding, Wei Li, +7 more
Organizations: Institute for AI Industry Research (AIR), Tsinghua University · HiThink Research · The Hong Kong University of Science and Technology (Guangzhou) · Zhejiang University · University of Science and Technology of China · SMBU · University of Washington · Mohamed bin Zayed University of Artificial Intelligence · Zhejiang University of Technology · Southeast University · Changan Automobile
Autonomous driving requires choosing a safe and efficient plan as surrounding traffic evolves. Generate-and-select planners propose multiple trajectories and score them for execution, and they have outperformed representative direct-prediction baselines on NAVSIM. Their scorer must compare plans that were never executed. Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained; a simulator, in contrast, can label the outcome of every candidate. We introduce World4Scorer, which builds the scorer as a trajectory-conditioned JEPA-style predictor: it predicts a state for each candidate and reads the candidate's scores from that state. Simulator outcome labels supervise the states of all candidates, and the observed future of the executed trajectory anchors the predictor to real scene evolution. Because one predictor produces every candidate's state, the anchor can constrain shared parameters used to score unexecuted plans, while the future itself is needed only during training. Generated candidates mostly score well, so a scene-matched bank adds low-scoring plans to the outcome supervision; framewise choices can conflict, so inertial re-ranking keeps consecutive selections consistent. World4Scorer achieves state-of-the-art NAVSIM-v2 performance and a strong adapted-system result on closed-loop Bench2Drive. With the LeWM world model and planning budget fixed, outcome-based scoring also improves manipulation planning on the OGBench-Cube benchmark.
Figures & tables
Figure 1: Gradient routing in generator–scorer planning. (a) Outcome labels train the scorer. (b) A world model adds future prediction, but only the executed trajectory has an observed future. (c) In World4Scorer, outcome labels supervise the state of every candidate, and the observed future anchors the predictor.
Figure 2: Overview of World4Scorer. (a) The trajectory generator proposes 64 candidates, scored by World4Scorer before inertial re-ranking. (b) The visual readout and score heads train the shared predictor: future observations supervise visual prediction for the executed trajectory; outcome labels supervise scoring for all generated candidates and 16 bank trajectories.
Table 1: Planning performance on NAVSIM-v1 and NAVSIM-v2 (%). Input C: cameras, L: LiDAR; superscripts 1 and 2 denote native NAVSIM-v1 and NAVSIM-v2. Best and second-best results are bold and underlined. We evaluate with the latest official NAVSIM devkits.