cs.CVFeb 20, 2026

VLANeXt: Recipes for Building Strong VLA Models

Authors: Xiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang, Runze Yang, Yihang Luo, Zhonghua Wu, Wei-Shi Zheng, +1 more

Organizations: 1S-Lab, Nanyang Technological University · 2ACE Robotics · 3Sun Yat-sen University · 4Shanghai Jiao Tong University · 5SenseTime Research

Abstract

Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving space, we reexamine the VLA design space under a unified framework and evaluation setup. Starting from a simple VLA baseline similar to RT-2, which is the origin of VLA, we systematically dissect design choices along three dimensions: foundational components, perception essentials, and action modelling perspectives. From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. The outcome of this exploration is a simple yet effective model, VLANeXt. It outperforms the state-of-the-art methods on the LIBERO and LIBERO-plus benchmarks and demonstrates strong performance in real-world experiments. We release a unified and easy-to-use codebase to reproduce our findings, explore the design space, and develop new VLA variants on top of a shared foundation. The codebase is available at https://github.com/DravenALG/VLANeXt.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models

    Mar 14, 2026Suhwan Choi, Yunsung Lee, Yubeen Park +4Preprocessing

  2. VLA Foundry: A Unified Framework for Training Vision-Language-Action Models

    Apr 21, 2026Jean Mercat, Sedrick Keh, Kushal Arora +5Vision-Language-Action FrameworkQwen3