A Survey on Efficient Vision-Language-Action Models
Organizations: School of Computer Science and Technology, Tongji University, China · School of Computing and Artificial Intelligence, Southwest Jiaotong University, China · School of Computer Science and Engineering, University of Electronic Science and Technology of China, China · Department of Information Engineering and Computer Science, University of Trento, Italy
Abstract
Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. To this end, recent studies improve VLA efficiency from different views, e.g., real-time inference, training computation, and scalable data collection. However, these efforts are mostly studied separately. A unified view is still missing for understanding how efficiency should be optimized across the full VLA lifecycle. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey provides an organized reference for the community and summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.
Figures & tables
| Models | Params.( ) | I.L.(ms) ( ) | Freq.(Hz) ( ) | Accelerator |
| RT-2-PaLI-X [ 3 ] | 55B | 330-1000 | 1-3 | Multi-TPU cloud |
| RT-2-PaLI-X [ 3 ] | 5B | 200 | 5 | Multi-TPU cloud |
| OpenVLA [ 1 ] | 7B | 166 | 6 | RTX 4090 |
| [ 2 ] | 3.3B | 73 | 20/50 | RTX 4090 |
| Hi Robot [ 58 ] | 3B | 73 | 10/50 | 1–2 RTX 4090 |
| GR00T N1 [ 12 ] | 2.2B | 63.9 | - | L40 GPU |
| Terms | Definition |
| Inference Latency | End-to-end time from input query to model output. |
| Control Frequency | The frequency at which a controller outputs control commands or action signals, typically measured in Hz. |
| Memory Footprint | Runtime or storage memory required by parameters, activations, and cached states. |
| Training Compute | Total computation consumed during training, typically measured by FLOPs or GPU-hours. |
| Data Efficiency | The data requirement for achieving target performance level; greater data efficiency corresponds to lower data requirements for comparable performance. |
| Deployment Efficiency | Practical efficiency under real deployment constraints, including latency, memory, and energy. |
| Survey | Year | Main Topic | Limitations |
| Ma et al. [ 23 ] | 2024 | General VLA taxonomy for embodied AI | Efficient VLAs scarcely discussed |
| Shao et al. [ 26 ] | 2025 | Large VLM-driven robotic manipulation VLAs | Only a few efficient fine-tuning methods mentioned |
| Xiang et al. [ 28 ] | 2025 | Human motor learning inspired VLA post-training | No efficient solution listed |
| Zhong et al. [ 27 ] | 2025 | Action tokenization methods in VLAs | Efficiency limited to tokenization dimension |
| Din et al. [ 24 ] | 2025 | Robotic manipulation oriented VLA review | Rarely discuss efficient solutions |
| Zhang et al. [ 25 ] | 2025 | Pure end-to-end VLA architectures | Only discuss part of efficient inference |
| Method/Model | Categories | Architecture(V/L/A) | Param. Scale | Key Innovations in Efficient Model Design |
| SARA-RT [ 69 ] | (A) | / / | ~5B(PaLI-X) | Quadratic-to-linear transformer up-training. |
| Long-VLA [ 70 ] | (A) | ResNet-18/CLIP/Diffusion Model(DDIM) | Phase-aware input masking, long-horizon attention optimization. | |
| RetoVLA [ 72 ] | (A) | SigLIP/SmolLM2/Flow Matching Model | Discarded register token reuse for spatial reasoning. | |
| KV-Efficient VLA [ 73 ] | (A) (I) | DINOv2+SigLIP/LLaMA-2/ | ~7B | RNN-gated chunked KV caching, selective context pruning. |
| dVLA [ 71 ] | (A) | MAGViT-v2/MMaDA/FAST | Prefix attention masking, KV caching for diffusion speedup. | |
| RoboMamba [ 19 ] | (B) (D) | CLIP/Mamba/MLP | ~3.2B | Linear-complexity Mamba reasoning, lightweight action decoding. |
| Method/Model | Training Phase | Strategies | Trainable modules | Key Innovations in Efficient Training |
| GeRM [ 93 ] | Pre-training | Off-RL | decoder-only Transformer + sparse MoE experts | Sparse MoE routing, offline RL over demonstration and sub-optimal data. |
| LAPA [ 43 ] | Pre-training | IL | VLM backbone + action quantization model | Unsupervised VQ-VAE quantization, label-free latent action extraction from video. |
| HAMSTER [ 98 ] | Pre-training | IL | vision encoder + LLM backbone + action module | Hierarchical action modeling, embodiment-agnostic 2D path representation. |
| DTP [ 42 ] | Multi-stage training | IL | diffusion trajectory model + action module | Diffusion-based trajectory synthesis, long-horizon task training configuration. |
| Humanoid-VLA [ 149 ] | Multi-stage training | IL | vision encoder + cross-attention module + LLM backbone + codebook | Language–motion pre-alignment, parameter-efficient egocentric visual adaptation. |
| GraspVLA [ 140 ] | Multi-stage training | IL | LLM backbone + projector + flow action expert | Synthetic data scaling, dual-stage grasping pre-training pipeline. |
| Method/Dataset | Category | Dataset Composition | Quality Control | Collection Acceleration |
| CLIP-RT [ 87 ] | (A) | 18.1M transitions,21K in-domain transitions | STA checks; heuristic labels. | Low-skill teleop + STA expansion. |
| GCENT [ 173 ] | (A) | 48K–62K real frames. | HDF5 validation; mode labels; Task Sentinel. | Failure rewind/refine; one-to-many supervision. |
| GeRM [ 93 ] | (B) | 257K simulation trajectories. | Success/failure labels; CQL penalty. | Automatic Isaac Gym rollouts. |
| GraspVLA [ 140 ] | (B) | 1B grasp frames. | CuRobo success; ray-traced randomization. | Parallel synthetic grasp generation. |
| cVLA [ 147 ] | (B) | 150k synthetic training samples. | Execution success; visibility/manual filters. | Randomized simulation pre-training. |
| RoboTwin 2.0 [ 16 ] | (B) | 100K+ expert trajectories. | 10 trials; log/VLM diagnosis; 5 repairs. | Code synthesis + simulation feedback. |
| Zhaoshu Yu is currently pursuing the Ph.D. degree in Computer Science and Technology at Tongji University, Shanghai, China. His research interests include vision-language-action models, embodied intelligence, and reinforcement learning. |
| Bo Wang is currently a junior undergraduate student majoring in Computer Science and Technology at Tongji University, Shanghai, China. His research interests include vision-language-action models, embodied intelligence, and multimodal learning. |
| Pengpeng Zeng received the B.E. degree from Xi’an University of Technology in 2016, and the M.E. and Ph.D. degrees from University of Electronic Science and Technology of China in 2019 and 2023, respectively. He is now a Researcher in Tongji University, China. His current research interests include visual understanding, machine learning, and reinforcement learning. |
| Haonan Zhang is currently a Postdoctoral Researcher with the School of Computer Science and Technology, Tongji University, China. He obtained his Ph.D. degree from the University of Electronic Science and Technology of China in 2026. His research interests include multimodal learning, intelligent dialogue, and agent-based systems. |
| Ji Zhang is an Assistant Professor with the School of Computing and Artificial Intelligence, Southwest Jiaotong University, China. He obtained his PhD degree from University of Electronic Science and Technology of China in 2024. His research interests include few-shot learning, transfer learning and robotics. |
| Zheng Wang received the B.E. and Ph.D. degrees from Zhejiang University, Hangzhou, China, in 2011 and 2017, respectively. He is currently with the School of Computer Science and Technology, Tongji University, Shanghai, China. His current research interests mainly focus on vision-language action, multimedia understanding, and computer vision. |
| Lianli Gao is a professor with the School of Computer Science and Engineering, University of Electronic Science and Technology of China. She obtained her PhD degree in Information Technology from The University of Queensland (UQ), Australia, under the supervision of Prof. Jane Hunter and Prof. Michael Bruenig. Her research ranges from Semantic Web, Machine Learning, Deep Learning, Computer Vision (Images and Videos), NLP, Knowledge Reasoning, Knowledge and the related practical applications etc. Specifically, she is mainly focusing on integrating Natural Language for Visual Content Understanding. She has the winner of the IEEE Transactions on Multimedia 2020 Prize Paper Award, Best Student Paper Award in Australian Database Conference (2017, Australia), IEEE TCMC Rising Star Award 2020 and ALIBABA Academic Young Fellow. She is an Associate Editor of IEEE TMM. |
| Jingkuan Song is a professor with the School of Computer Science and Technology, Tongji University, China. He joined Columbia University as a Postdoc Research Scientist (2016-2017), and University of Trento as a Research Fellow (2014-2016). He obtained his PhD degree in 2014 from The University of Queensland (UQ), Australia. His research interest includes large-scale multimedia retrieval, LLMs and deep learning techniques. He was the winner of the Best Paper Award in ICPR (2016, Mexico), Best Student Paper Award in Australian Database Conference (2017, Australia), and Best Paper Honorable Mention Award (2017, Japan). He is an Associate Editor of IEEE TMM and ACM TOMM. |
| Nicu Sebe is a professor with the Department of Information Engineering and Computer Science, University of Trento, leading the research in the areas of multimedia information retrieval and human behavior understanding. He was the General CoChair of ACM Multimedia 2013 and 2022, and the Program Chair of ACM Multimedia 2007 and 2011, ECCV 2016, ICCV 2017 and ICPR 2020. He is a fellow of the International Association for Pattern Recognition (IAPr) and of the European Laboratory for Learning and Intelligent Systems (ELLIS). |
| Heng Tao Shen is a professor with the School of Computer Science and Technology, Tongji University, China. He obtained his BSc with 1st class Honours and PhD from Department of Computer Science, National University of Singapore in 2000 and 2004 respectively. His current research interests include multimedia search, computer vision, artificial intelligence, and big data management. He has published 300+ peer-reviewed papers and received 7 best paper awards from international conferences, including the Best Paper Award from ACM Multimedia 2017 and Best Paper Award-Honourable Mention from ACM SIGIR 2017. He has served as General Co-chair for ACM Multimedia 2021 and TPC Co-Chair for ACM Multimedia 2015, and is an Associate Editor of ACM Trans. of Data Science (TDS), IEEE Trans. on Image Processing (TIP), IEEE Trans. on Multimedia (TMM), and IEEE Trans. on Knowledge and Data Engineering (TKDE). He is a Fellow of ACM/IEEE/OSA. |