cs.ROSep 26, 2026

Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation

Authors: Xuekang Yang, Lu Chen, Shuang Luo, Jialing Zhu, Qi Zhang, Yue Gao, Xiang Zhang

Organizations: School of Computer Science, Shanghai Jiao Tong University · School of Automation and Intelligent Sensing, Shanghai Jiao Tong University · Defense Innovation Institute, Academy of Military Sciences · MoE Key Laboratory of Artificial Intelligence and AI Institute, Shanghai Jiao Tong University

Abstract

Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify an executable affordance pose and recover from accumulated action errors. We propose PACE (Preference-refined Affordance-Conditioned Execution), a supervised local execution module that augments frozen zero-shot semantic planners for reliable cross-floor navigation. PACE grounds transition-related semantics into a long-horizon, agent-centric traversable affordance pose and conditions short-horizon action generation on this spatial target, thereby aligning semantic goals with physical execution. We further post-train PACE through failure-aware preference refinement using rollout-derived pairs that contrast normal or recovery behaviors with deviation-amplifying behaviors, thereby improving closed-loop correction. We integrate PACE into six open-source zero-shot VLN navigators and demonstrate consistent improvements on the cross-floor subsets of R2R-CE and RxR-CE, increasing the average success rate from 16.35% to 27.65% and from 4.76% to 12.06%, respectively. Real-world experiments further demonstrate PACE's applicability in unseen environments, highlighting the potential of traversable affordances to bridge semantic intent and reliable embodied behavior.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TravExplorer: Cross-Floor Embodied Exploration via Traversability-Aware 3-D Planning

    May 19, 2026Han Zheng, Zhe Chen, Yudong Huang +4Dynamic EnvironmentsExploration

  2. SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning

    Jun 8, 2026Yucheng Deng, Pingrui Lai, Xinhai Li +5Vision-Language NavigationSpatial Memory

  3. P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation

    May 19, 2026Kai Sheng, Liuyi Wang, Haojie Dai +5Vision-Language NavigationMenu