cs.CVOct 27, 2025

A Survey on Efficient Vision-Language-Action Models

Authors: Zhaoshu Yu, Bo Wang, Pengpeng Zeng, Haonan Zhang, Ji Zhang, Zheng Wang, Lianli Gao, Jingkuan Song, +2 more

Organizations: School of Computer Science and Technology, Tongji University, China · School of Computing and Artificial Intelligence, Southwest Jiaotong University, China · School of Computer Science and Engineering, University of Electronic Science and Technology of China, China · Department of Information Engineering and Computer Science, University of Trento, Italy

Abstract

Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. To this end, recent studies improve VLA efficiency from different views, e.g., real-time inference, training computation, and scalable data collection. However, these efforts are mostly studied separately. A unified view is still missing for understanding how efficiency should be optimized across the full VLA lifecycle. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey provides an organized reference for the community and summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.

Figures & tables

Explore similar work

CardsList
  1. VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

    Jul 7, 2025Juyi Lin, Amir Taherin, Arash Akbari +11Diffusion-Based Vision-Language-ActionsRobotic Manipulation

  2. Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines

    Apr 24, 2026Ziyao Wang, Bingying Wang, Hanrong Zhang +7EmbodimentAction Space

  3. Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

    Jul 14, 2026Yuzhou Wu, Yuxin Zheng, Muchun Niu +6Diffusion-Based Vision-Language-ActionsVision-Language-Action Framework