Organizations: School of Computer Science, Shanghai Jiao Tong University, Shanghai, P. R. China · College of Information Science and Technology, Eastern Institute of Technology, Ningbo, P. R. China · School of Mechanical and Power Engineering, Zhengzhou university, Zhengzhou,P.R.China · University of Waterloo
Autonomous navigation in cluttered and constrained environments typically assumes a fixed environment and searches only for paths within existing free space. However, a route can be blocked by an articulated structure, a movable object, or a pedestrian, and reaching the goal may therefore require appropriate embodied interaction with the environment. This paper formulates Path-Creative Navigation (PCN), a navigation paradigm in which the robot recovers free space by coordinating locomotion and embodied interaction, and addresses one class of PCN tasks with a mapless framework. The vision-LiDAR-odometry framework integrates local observations, obstacle geometry, and robot-state estimates through a unified navigation system, providing stable semantic and geometric information for navigation and interaction. The traversability-aware decision method utilizes visual and LiDAR distance information to determine whether a blockage is actionable and whether it can be safely bypassed, enabling the robot to interact only when necessary while supporting articulated structure pushing, movable object pushing, obstacle avoidance, and pedestrian requests. We design simulation scenarios and evaluate different methods in both simulation and two real-world tasks that cannot be completed by conventional navigation alone without embodied interaction. These results demonstrate that our framework achieves higher task completion than the baseline methods by coordinating navigation and interaction in cluttered and constrained environments. Code is available at:{https://anonymous.4open.science/r/path-creative-navigation/}.
Figures & tables
Fig. 2: Overview of the framework. Vision, depth, LiDAR, and odometry are integrated to obtain stable object targets ot,i and local geometric observations for closed-loop interaction. The traversability-aware decision method jointly evaluates object actionability, Corridor(ot,i) , and NoBypass(ot,i) to determine NeedInteract(ot,i) . The robot therefore retains navigation for fixed or bypassable obstacles and invokes the corresponding embodied interaction only when necessary. After the scene changes, traversability is reassessed and navigation resumes toward the goal.
Fig. 3: Objects in simulation. The simulated environments include doors of different sizes, colors, and types; boxes of varying colors and sizes; and chairs with different sizes and orientations.
Fig. 4: Experimental trajectories in simulation. Our method successfully performs the necessary interactions while treating non-blocking actionable objects as normal obstacles, avoiding unnecessary interactions and reaching the goal.
Table 4
Fig. 5: Average execution time (left) and trial-level pass rate (right) of the evaluated methods.
Fig. 6: Real-world embodied interaction. Interaction executions include articulated structure push and movable obstacle push. The pedestrian request action makes the robot play a verbal request and stop moving until NeedInteract(ot,i)=0 .
Robot navigation typically assumes an obstacle-free path exists between start and goal. In real environments, however, clutter may block all routes. We introduce Lifelong Interactive Navigation, where a mobile robot with manipulation capabilities must move objects to forge paths and complete sequential object-placement tasks. Because environment modifications persist, decisions impact future navigability and task difficulty. We propose CoReLIN, an LLM-driven constraint-based reasoning framework with active perception. CoReLIN reasons over a structured scene graph to decide which objects to relocate, where to place them, and where to explore next. A standard motion planner executes reliable navigation and manipulation primitives. To evaluate long-horizon behavior, we introduce 2 new metrics - Long-term Efficiency Score (LES), a unified metric capturing success, execution efficiency, environment optimality, captured by Price of Clutter. In ProcTHOR-10k, CoReLIN outperforms best baseline by 16% under standard metrics and LES, and transfers to real-world hardware.
Goal-conditioned navigation models for ground robots trained using supervised learning show promising zero-shot transfer, but their collision-avoidance capability nevertheless degrades under distribution shift, i.e. environmental, robot or sensor configuration changes. We propose ViLiNT a multimodal, attention-based policy for goal navigation, trained on heterogeneous data from multiple platforms and environments, which improves robustness with two key features. First, we fuse RGB images, 3D LiDAR point clouds, a goal embedding and a robot's embodiment descriptor with a transformer architecture to capture complementary geometry and appearance cues. The transformer's output is used to condition a diffusion model that generates navigable trajectories. Second, using automatically generated offline labels, we train a path clearance prediction head for scoring and ranking trajectories produced by the diffusion model. The diffusion conditioning as well as the trajectory ranking head depend on a robot's embodiment token that allows our model to generate and select trajectories with respect to the robot's dimensions. Across three simulated environments, ViLiNT improves Success Rate on average by 166% over equivalent state-of-the-art vision-only baseline (NoMaD). This increase in performance is confirmed through real-world deployments of a rover navigating in obstacle fields. These results highlight that combining multimodal fusion with our collision prediction mechanism leads to improved off-road navigation robustness.
Louis Dezons, Quentin Picard, Rémi Marsal +2
AMIAD Pˆole Recherche, Palaiseau, France · U2IS, ENSTA, Institut Polytechnique de Paris, Palaiseau, France
Robotic manipulators operating in unstructured environments face significant challenges in safely executing goal-directed tasks due to dynamic and unforeseen obstacles, while traditional methods rely on prior knowledge or fixed perception pipelines, limiting adaptability. We propose a framework for safe task execution with effective obstacle avoidance. The environment module performs real-time obstacle detection, 3D localization, and ground surface geometry estimation. It then generates a structured semantic report that includes obstacle positions, object geometry and shape, and whether obstacles lie inside, outside, or within critical interaction zones. A central coordination module manages the overall system by handling tool invocation (e.g., memory and MoveIt collision scene updates), facilitating communication between modules, and continuously monitoring task progress until completion. Furthermore, a planning module selects an appropriate motion planning algorithm, such as RRTConnect, RRT*, or BiTRRT, based on the current environment configuration and goal requirements. The trajectory generated by the planner is further analyzed and refined to ensure safe and collision-free task execution. The proposed approach is evaluated in Gazebo Classic , demonstrating robustness in dynamic scenarios.
Aachal Sharma, Narendra Kumar Dhar
Centre for Artificial Intelligence and Robotics (CAIR), Indian Institute of Technology Mandi, Himachal Pradesh 175005, India.