cs.RODate pending

NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models

Authors: Ziyue ZhuShangyang WuShuai ZhaoZhiqiu ZhaoJian ZhangShengjie LiYi WangAnh Tuan Luu+3 more

Organizations: 1Beijing University of Posts and Telecommunications · 2Nanyang Technological University

Abstract

Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face structure-blind backbones, backbone-bound generalization, and flat single-objective optimization. To address these challenges, we propose a novel Neuro-Symbolic Vision-Language-Action (NS-VLA) framework. It introduces a Neuro-Symbolic Encoder for plan-constrained primitive inference, a Neuro-Symbolic Solver that conditions a backbone-agnostic policy on the active primitive, and Hierarchical Joint Policy Optimization with reward-granularity matching. Experiments on robotic manipulation benchmarks demonstrate that NS-VLA outperforms previous methods in both one-shot training and data-perturbed settings, while simultaneously exhibiting superior zero-shot generalizability and expanded exploration space. Our code is publicly available.