A Survey on Efficient Vision-Language-Action Models
Authors: Zhaoshu Yu, Bo Wang, Pengpeng Zeng, Haonan Zhang, Ji Zhang, Zheng Wang, Lianli Gao, Jingkuan Song, +2 more
Organizations: School of Computer Science and Technology, Tongji University, China · School of Computing and Artificial Intelligence, Southwest Jiaotong University, China · School of Computer Science and Engineering, University of Electronic Science and Technology of China, China · Department of Information Engineering and Computer Science, University of Trento, Italy
Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. To this end, recent studies improve VLA efficiency from different views, e.g., real-time inference, training computation, and scalable data collection. However, these efforts are mostly studied separately. A unified view is still missing for understanding how efficiency should be optimized across the full VLA lifecycle. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey provides an organized reference for the community and summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.
Figures & tables
Fig. 1: The transition from foundational VLAs to efficient VLAs. Foundational VLAs are bottlenecked by real-time incompatibility, excessive computational costs, and inefficient data collection. To enable deployment on edge devices, Efficient VLAs navigate a design space characterized by diverse trade-offs across three core pillars. By strategically balancing these competing objectives, different methods achieve low latency, rapid adaptation, and high reliability across diverse edge device implementations.
Fig. 2: The organization of our survey. We systematically categorize efficient VLAs into three core pillars: (1) Efficient Model Design , encompassing efficient architectures and model compression techniques, (2) Efficient Training , covering efficient pre-training and post-training strategies, and (3) Efficient Data Collection , including efficient data collection and augmentation methods. The framework also reviews VLA foundations and key applications.
Fig. 3: An overview of VLAs. VLAs integrate vision encoders to extract visual features, LLM backbones to fuse multimodal inputs, and action decoders (autoregressive or generative) to produce robotic control signals, enabling end-to-end vision-language-action reasoning for embodied manipulation tasks.
Fig. 4: Timeline of foundational and efficient VLAs. This timeline charts the progression toward resource-conscious embodied intelligence. Inclusion in the “Efficient VLA” trajectory strictly requires explicit contributions addressing foundational bottlenecks via at least one of our core taxonomy pillars: mitigating computational footprint and inference latency (Efficient Model Design), minimizing adaptation overhead (Efficient Training), or reducing data acquisition costs (Efficient Data Collection).
Models
Params.( ↓ )
I.L.(ms) ( ↓ )
Freq.(Hz) ( ↑ )
Accelerator
RT-2-PaLI-X [ 3 ]
55B
330-1000
1-3
Multi-TPU cloud
RT-2-PaLI-X [ 3 ]
5B
200
5
Multi-TPU cloud
OpenVLA [ 1 ]
7B
166
6
RTX 4090
π0 [ 2 ]
3.3B
73
20/50
RTX 4090
Hi Robot [ 58 ]
3B
73
10/50
1–2 RTX 4090
GR00T N1 [ 12 ]
2.2B
63.9
-
L40 GPU
TABLE I: Efficiency-related metrics of representative VLAs. The table reports the number of parameters ( Params. ), inference latency ( I.L. ), operating frequency ( Freq. ), and accelerator ( Accelerator ) of representative VLAs to contextualize cross-paper efficiency comparisons. “ ↓ ” indicates that lower values are better, and “ ↑ ” indicates that higher values are better. “/” denotes that the model can operate at two different frequencies.
Terms
Definition
Inference Latency
End-to-end time from input query to model output.
Control Frequency
The frequency at which a controller outputs control commands or action signals, typically measured in Hz.
Memory Footprint
Runtime or storage memory required by parameters, activations, and cached states.
Training Compute
Total computation consumed during training, typically measured by FLOPs or GPU-hours.
Data Efficiency
The data requirement for achieving target performance level; greater data efficiency corresponds to lower data requirements for comparable performance.
Deployment Efficiency
Practical efficiency under real deployment constraints, including latency, memory, and energy.
TABLE II: Glossary of VLA efficiency metrics. Standardized definitions of core efficiency dimensions across the model-training-data lifecycle.
Survey
Year
Main Topic
Limitations
Ma et al. [ 23 ]
2024
General VLA taxonomy for embodied AI
Efficient VLAs scarcely discussed
Shao et al. [ 26 ]
2025
Large VLM-driven robotic manipulation VLAs
Only a few efficient fine-tuning methods mentioned
Xiang et al. [ 28 ]
2025
Human motor learning inspired VLA post-training
No efficient solution listed
Zhong et al. [ 27 ]
2025
Action tokenization methods in VLAs
Efficiency limited to tokenization dimension
Din et al. [ 24 ]
2025
Robotic manipulation oriented VLA review
Rarely discuss efficient solutions
Zhang et al. [ 25 ]
2025
Pure end-to-end VLA architectures
Only discuss part of efficient inference
TABLE III: Comparison of related VLA surveys. Prior works predominantly focus on general taxonomies or specific mechanisms, leaving a noticeable gap in systematic efficiency analysis.
Fig. 5: Key strategies for Efficient Architectures (Sec. III-A ), including Efficient Attention, Transformer Alternatives, Efficient Action Decoding, Lightweight Components, Mixture-of-Experts, and Hierarchical Systems.
TABLE IV: Representative works on Efficient Model Design. "Categories" markers denote subcategories in Sec. III : (A) Efficient Attention, (B) Transformer Alternatives, (C) Efficient Action Decoding, (D) Lightweight Component, (E) Mixture-of-Experts, (F) Hierarchical Systems, (G) Layer Pruning, (H) Quantization, and (I) Token Optimization. For general methods, the smallest experimental parameter scale is reported, with unavailable data denoted by “ − ”.
Fig. 6: Key strategies for Model Compression (Sec. III-B ), including Layer Pruning, Quantization, and Token Optimization through compression, pruning, and caching.
Fig. 7: Key strategies for Efficient Training (Sec. IV ), comprising Efficient Pre-Training and Efficient Post-Training with their corresponding strategies.
Method/Model
Training Phase
Strategies
Trainable modules
Key Innovations in Efficient Training
GeRM [ 93 ]
Pre-training
Off-RL
decoder-only Transformer + sparse MoE experts
Sparse MoE routing, offline RL over demonstration and sub-optimal data.
LAPA [ 43 ]
Pre-training
IL
VLM backbone + action quantization model
Unsupervised VQ-VAE quantization, label-free latent action extraction from video.
Synthetic data scaling, dual-stage grasping pre-training pipeline.
TABLE V: Representative works on Efficient Training. The “Training Phase” column denotes the specific learning stage where efficient optimizations are targeted. The “Strategies” markers “IL”, “On-RL”, and “Off-RL” correspond to efficiency-focused Imitation Learning, Online Reinforcement Learning, and Offline Reinforcement Learning paradigms, respectively. “Trainable modules” column outlines the specific architectural components updated during training.
Fig. 8: Taxonomy of Efficient Data Collection strategies in VLAs. This figure illustrates the primary approaches under Sec. V , encompassing human-in-the-loop, simulated, reusability-oriented, self-driven, and augmentative techniques aimed at scaling dataset acquisition and improving data collection efficiency.
Method/Dataset
Category
Dataset Composition
Quality Control
Collection Acceleration
CLIP-RT [ 87 ]
(A)
18.1M transitions,21K in-domain transitions
STA checks; heuristic labels.
Low-skill teleop + STA expansion.
GCENT [ 173 ]
(A)
48K–62K real frames.
HDF5 validation; mode labels; Task Sentinel.
Failure rewind/refine; one-to-many supervision.
GeRM [ 93 ]
(B)
257K simulation trajectories.
Success/failure labels; CQL penalty.
Automatic Isaac Gym rollouts.
GraspVLA [ 140 ]
(B)
1B grasp frames.
CuRobo success; ray-traced randomization.
Parallel synthetic grasp generation.
cVLA [ 147 ]
(B)
150k synthetic training samples.
Execution success; visibility/manual filters.
Randomized simulation pre-training.
RoboTwin 2.0 [ 16 ]
(B)
100K+ expert trajectories.
10 trials; log/VLM diagnosis; ≤ 5 repairs.
Code synthesis + simulation feedback.
TABLE VI: Representative works on Efficient Data Collection. Categories follow Sec. V : (A) Human-in-the-loop Data Collection, (B) Simulation Data Collection, (C) Internet-Scale and Cross-Domain Data Utilization, (D) Self-Exploration Data Collection, and (E) Data Augmentation. “Dataset Composition” reports quantity only, and “ − ” means not specified. “Quality Control” summarizes data validation or correction. “Collection” Acceleration summarizes reductions in real-robot data collection.
Zhaoshu Yu is currently pursuing the Ph.D. degree in Computer Science and Technology at Tongji University, Shanghai, China. His research interests include vision-language-action models, embodied intelligence, and reinforcement learning.
Table 15
Bo Wang is currently a junior undergraduate student majoring in Computer Science and Technology at Tongji University, Shanghai, China. His research interests include vision-language-action models, embodied intelligence, and multimodal learning.
Table 16
Pengpeng Zeng received the B.E. degree from Xi’an University of Technology in 2016, and the M.E. and Ph.D. degrees from University of Electronic Science and Technology of China in 2019 and 2023, respectively. He is now a Researcher in Tongji University, China. His current research interests include visual understanding, machine learning, and reinforcement learning.
Table 17
Haonan Zhang is currently a Postdoctoral Researcher with the School of Computer Science and Technology, Tongji University, China. He obtained his Ph.D. degree from the University of Electronic Science and Technology of China in 2026. His research interests include multimodal learning, intelligent dialogue, and agent-based systems.
Table 18
Ji Zhang is an Assistant Professor with the School of Computing and Artificial Intelligence, Southwest Jiaotong University, China. He obtained his PhD degree from University of Electronic Science and Technology of China in 2024. His research interests include few-shot learning, transfer learning and robotics.
Table 19
Zheng Wang received the B.E. and Ph.D. degrees from Zhejiang University, Hangzhou, China, in 2011 and 2017, respectively. He is currently with the School of Computer Science and Technology, Tongji University, Shanghai, China. His current research interests mainly focus on vision-language action, multimedia understanding, and computer vision.
Table 20
Lianli Gao is a professor with the School of Computer Science and Engineering, University of Electronic Science and Technology of China. She obtained her PhD degree in Information Technology from The University of Queensland (UQ), Australia, under the supervision of Prof. Jane Hunter and Prof. Michael Bruenig. Her research ranges from Semantic Web, Machine Learning, Deep Learning, Computer Vision (Images and Videos), NLP, Knowledge Reasoning, Knowledge and the related practical applications etc. Specifically, she is mainly focusing on integrating Natural Language for Visual Content Understanding. She has the winner of the IEEE Transactions on Multimedia 2020 Prize Paper Award, Best Student Paper Award in Australian Database Conference (2017, Australia), IEEE TCMC Rising Star Award 2020 and ALIBABA Academic Young Fellow. She is an Associate Editor of IEEE TMM.
Table 21
Jingkuan Song is a professor with the School of Computer Science and Technology, Tongji University, China. He joined Columbia University as a Postdoc Research Scientist (2016-2017), and University of Trento as a Research Fellow (2014-2016). He obtained his PhD degree in 2014 from The University of Queensland (UQ), Australia. His research interest includes large-scale multimedia retrieval, LLMs and deep learning techniques. He was the winner of the Best Paper Award in ICPR (2016, Mexico), Best Student Paper Award in Australian Database Conference (2017, Australia), and Best Paper Honorable Mention Award (2017, Japan). He is an Associate Editor of IEEE TMM and ACM TOMM.
Table 22
Nicu Sebe is a professor with the Department of Information Engineering and Computer Science, University of Trento, leading the research in the areas of multimedia information retrieval and human behavior understanding. He was the General CoChair of ACM Multimedia 2013 and 2022, and the Program Chair of ACM Multimedia 2007 and 2011, ECCV 2016, ICCV 2017 and ICPR 2020. He is a fellow of the International Association for Pattern Recognition (IAPr) and of the European Laboratory for Learning and Intelligent Systems (ELLIS).
Table 23
Heng Tao Shen is a professor with the School of Computer Science and Technology, Tongji University, China. He obtained his BSc with 1st class Honours and PhD from Department of Computer Science, National University of Singapore in 2000 and 2004 respectively. His current research interests include multimedia search, computer vision, artificial intelligence, and big data management. He has published 300+ peer-reviewed papers and received 7 best paper awards from international conferences, including the Best Paper Award from ACM Multimedia 2017 and Best Paper Award-Honourable Mention from ACM SIGIR 2017. He has served as General Co-chair for ACM Multimedia 2021 and TPC Co-Chair for ACM Multimedia 2015, and is an Associate Editor of ACM Trans. of Data Science (TDS), IEEE Trans. on Image Processing (TIP), IEEE Trans. on Multimedia (TMM), and IEEE Trans. on Knowledge and Data Engineering (TKDE). He is a Fellow of ACM/IEEE/OSA.
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of massive tokens leading to high inference latency and increased training cost, and (ii) insufficient utilization of generated actions resulting in potential performance loss. To address these issues, we develop a training framework to finetune VLA models for generating significantly fewer action tokens with high parallelism, effectively reducing inference latency and training cost. Furthermore, we introduce an inference optimization technique with a novel voting-based ensemble strategy to combine current and previous action predictions, improving the utilization of generated actions and overall performance. Our results demonstrate that we achieve superior performance compared with state-of-the-art VLA models, achieving significantly higher success rates and 39× faster inference than OpenVLA with 46 Hz throughput on edge platforms, demonstrating practical deployability. The code is available at https://github.com/LukeLIN-web/VOTE.
Juyi Lin, Amir Taherin, Arash Akbari +11
Northeastern University, Boston, USA · EmbodyX,San Mateo,USA
Despite remarkable progress in Vision--Language--Action (VLA) models, a central bottleneck remains underexamined: the data infrastructure that underlies embodied learning. In this survey, we argue that future advances in VLA will depend less on model architecture and more on the co-design of high-fidelity data engines and structured evaluation protocols. To this end, we present a systematic, data-centric analysis of VLA research organized around three pillars: datasets, benchmarks, and data engines. For datasets, we categorize real-world and synthetic corpora along embodiment diversity, modality composition, and action space formulation, revealing a persistent fidelity-cost trade-off that fundamentally constrains large-scale collection. For benchmarks, we analyze task complexity and environment structure jointly, exposing structural gaps in compositional generalization and long-horizon reasoning evaluation that existing protocols fail to address. For data engines, we examine simulation-based, video-reconstruction, and automated task-generation paradigms, identifying their shared limitations in physical grounding and sim-to-real transfer. Synthesizing these analyses, we distill four open challenges: representation alignment, multimodal supervision, reasoning assessment, and scalable data generation. Addressing them, we argue, requires treating data infrastructure as a first-class research problem rather than a background concern.
Ziyao Wang, Bingying Wang, Hanrong Zhang +7
University of Maryland, College Park · University of Utah · Northeastern University +1
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2