Safe navigation in dynamic environments requires anticipating future environmental states to account for spatiotemporal risks, specifically when and where collisions may occur. To this end, occupancy grid map (OGM) prediction has been widely adopted as an effective approach. However, existing OGM-based navigation methods often struggle to achieve accurate and efficient forecasting and fail to fully exploit the temporal information in predicted OGMs during planning. To address these challenges, we propose FORTE, a navigation framework that directly exploits the spatiotemporal evolution of predicted occupancy from the perspectives of spatiotemporal occupancy overlap and occupancy directivity. Based on these properties, FORTE evaluates multiple topology-distinct paths and selects the suitable one without explicit object detection or tracking. To support online planning, we formulate a latent diffusion model-based OGM predictor that generates the entire forecast horizon in a non-autoregressive manner while maintaining temporal consistency through temporal shift modules. Extensive evaluations demonstrate that FORTE outperforms state-of-the-art baselines. For prediction, FORTE achieves up to 215.3% higher IoU and 5.24x faster inference; for navigation, it yields up to a 3.5x higher success rate.
Figures & tables
Fig. 1 : FORTE Pipeline. (a–c) OGM forecasting: At inference, recent LiDAR scans and robot states are used to forecast OGMs using LDM, aligned with the current frame, yielding o~t+1:t+τ for planning. (d–g) TSR planning: Given the current OGM ot and a reference path πref , the planner generates topology-distinct candidate guidance paths and evaluates their arrival-time and directivity costs using the predicted occupancy evolution. The selected path π∗ guides the local planner.
Fig. 2 : Directivity cost from predicted occupancy propagation. (a) A candidate guidance path and an approaching obstacle. (b) The path points are mapped into the forecasted OGM sequence using estimated arrival times. (c) Occupancy peaks at adjacent path points pj and pj+1 are connected to form an edge e∈Ti . (d) Occupancy peaks earlier at the farther point pj+1 ( m′=2 ) than at the nearer point pj ( m=3 ). (e) m′<m indicates that occupancy propagates toward the robot along the path, thereby incurring a positive directivity cost; the reverse ordering is shown in the dashed box.
Fig. 3 : Prediction performance. Average IoU, SSIM, and NMI at each of the 10 prediction steps, averaged over the test set. The curves show the mean over 8 stochastic prediction samples.
Fig. 4 : Simulation Scenarios. Six navigation scenarios with diverse pedestrian motions and spatial constraints. Blue arrows indicate the pedestrian motion directions. Start positions (yellow circles) and goal positions (red crosses) are marked.
Local Planner
Method
Scenario 1
Scenario 2
Scenario 3
Scenario 4
Scenario 5
Scenario 6
SR (%)
TTG (s)
SR (%)
TTG (s)
SR (%)
TTG (s)
SR (%)
TTG (s)
SR (%)
TTG (s)
SR (%)
TTG (s)
DWA
SCOPE
14
26.1
56
22.1
16
30.55
18
21.63
22
49.32
32
31.4
SCOPE++
30
28.0
72
21.71
6
30.13
60
21.82
8
56.52
2
33
SO-SCOPE
22
30.3
60
22.8
6
30.03
48
22.92
0
–
12
31.2
FORTE w/o TSR planner
44
22.86
74
20.36
58
25.87
66
20.24
22
52.78
72
29.74
FORTE w/o Directivity cost
82
19.4
84
16.72
68
19.21
68
18.12
30
42.1
76
23.87
TABLE I : Navigation performance in Simulation. Success Rate (SR) and Time to Goal (TTG) are reported for each scenario; TTG is for successful trials only.
Fig. 5 : Real-world Scenarios. (a) Two pedestrians moving in opposite directions. (b) Multiple pedestrians moving in various directions. Both scenarios are set in corridor environments. Red curves indicate the robot paths.
Future 3D semantic occupancy forecasting and motion planning are central to autonomous driving, as they require models to reason about how surrounding scenes evolve and how the ego vehicle should act. Existing occupancy world models commonly discretize scenes into latent embeddings, volumetric features, or quantized tokens, and forecast future states through fixed-step autoregressive generation. This limits temporal flexibility, obscures scene evolution, accumulates errors over long horizons, and poorly matches the continuous-time dynamics of real driving scenes. We propose GEM, a Gaussian Evolution Model for non-autoregressive occupancy world modeling, where driving scenes are represented as explicit continuous 4D Gaussian primitives with learned dynamics. Instead of rolling out future occupancy states step by step, GEM directly queries the Gaussian world representation at arbitrary timestamps and splats the corresponding conditional 3D Gaussians into semantic occupancy volumes. This enables efficient forecasting over the full horizon while retaining a compact and interpretable scene representation. By decoupling spatial geometry, temporal support, and primitive motion, GEM makes the predicted world easier to inspect, as each primitive's evolution can be followed continuously over time. The same representation also supports motion planning by predicting future ego trajectories from the learned Gaussian world. Extensive experiments show that GEM achieves state-of-the-art future semantic occupancy forecasting and strong motion planning performance, while providing flexible temporal querying.
Legged robots under sparse waypoint guidance must avoid moving obstacles using partial, rapidly changing LiDAR observations. We present LOOP (Latent-recurrent Occupancy rollOut Policy), a local avoidance policy that connects sparse waypoint guidance to a frozen locomotion controller at 50 Hz. From occupancy and ego-velocity histories, a recurrent predictor forecasts future occupancy over a 1 s horizon by warping the current map with learned flow and visibility gates. These maps guide velocity selection through map-derived features and geometric risk estimates, providing an explicit interface for inspecting and replacing predictions. In encounter-synchronised Isaac Lab evaluations, LOOP achieves 57.1% head-on success at obstacle speeds of 2.5-3.2 m/s, exceeding a retrained reactive baseline by 8.2 percentage points. Comparisons with a rollout-free BEV policy show smaller, scenario-dependent gains from the prediction branch, including improved crossing success and reduced variability across training seeds at the highest head-on speeds. The adapter runs onboard a Unitree Go2 in 14.5 ms per step and completes all 16 real-world crossing trials without collision, demonstrating deployment feasibility.
Yuhui Mao, Fen Liu, Shenghai Yuan +3
School of Electrical and Electronic Engineering, Nanyang Technological University, 50 Nanyang Avenue, Singapore 639798
Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong spatial-temporal modeling between the input multi-view frames is still underexplored, which limits the performance of those methods in future 4D forecasting. To address this gap, we introduce a novel framework, InterOCF, for 4D occupancy forecasting that jointly models temporal dynamics in both 3D voxel-based representations and multi-view segmentation sequences, while explicitly incorporating feature interaction between the 2D and 3D branches. Our framework incorporates three core components: 1) A 3D Spatio-Temporal (3DST) module that learns volumetric dynamics from historical voxel states to predict future voxel states; 2) A 2D Spatio-Temporal (2DST) module employing an auxiliary multi-view temporal segmentation forecasting task to enhance temporal semantic dynamics; 3) A Spatio-Temporal Interaction Modeling (STIM) module that enables feature interaction between 2D and 3D representations. Experiments on the nuScenes, Lyft-Level5, and nuScenes-Occupancy datasets show that InterOCF consistently outperforms existing baseline approaches.
Qi Zhang, Xinquan Yu, Kaiyi Zhang +1
College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China.