cs.CVOct 8, 2026

VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation

Authors: Boyao Han, Chen Shi, Jingjing Qian, ZhuoTan Tian, Li Jiang

Organizations: The Chinese University of Hong Kong, Shenzhen · Shenzhen Loop Area Institute · Harbin Institute of Technology, Shenzhen

Abstract

Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained base VLA as a supplementary visual condition, requiring no explicit 3D sensing or novel-view rendering. Experiments on RoboTwin, LIBERO, and a real-robot platform demonstrate that VersaCamVLA consistently outperforms prior VLA methods and direct multi-view baselines, maintaining robust performance across varying camera counts and unseen camera poses.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models

    Oct 8, 2026Junmyeong Lee, Dongmin Shin, Min-Gyu Park +3Robustness of VLA ModelsVision-Language-Action Models

  2. From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model

    Jul 6, 2026Wenhao Li, Xueying Jiang, Quanhao Qian +4Language-Conditioned Robot ManipulationHand-Eye Calibration

  3. G3^3VLA: Geometric inductive bias for Vision-Language-Action Models

    Jun 23, 2026Yue Peng, Yongzhe Zhao, Artur Habuda +5Multi-View LearningRobotic Manipulation