cs.LGMar 9, 2026

DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models

Authors: Zihao Zheng, Hangyu Cao, Chibang Tao, Zhihao Mao, Lingyue Zhang, Maoliang Li, Jiayu Chen, Hailong Zou, +4 more

Organizations: School of Computer Science, Peking University, Beijing, China. · School of Software Engineering, South China University of Technology, Guangzhou, China. · School of Artificial Intelligence, Beijing Normal University, Beijing, China. · School of Electronics Engineering and Computer Science, Peking University, Beijing China.

Abstract

Vision-language-action (VLA) models achieve favorable task performance, yet runtime errors in closed-loop execution evolve with alternating updates of actions and observations. Existing VLA quantization methods mainly trade off inference speed and model performance while rarely investigating how quantization alters runtime errors. By comparing error distributions between full-precision and quantized models, we identify distinct evolutionary patterns for the two types of errors: quantization enlarges the variance of translation error distributions and increases their dispersion, whereas the distribution center of rotation errors gradually shifts across execution steps, demonstrating cumulative drift. Motivated by this observation, we rethink the optimal quantization strategy for VLA models and propose \textit{DyQ-VLA}, a runtime-error-aware quantization framework, which dynamically selects activation precision according to execution steps and tracks as well as compensates rotation errors via accumulated quantization residuals. Corresponding operators and runtime adaptation strategies are devised within the framework to enable dynamic-precision execution. Experiments show that \textit{DyQ-VLA} achieves a 1.88 to 1.93 inference speedups while maintaining comparable or higher average task success rates. Moreover, it reduces mean execution steps by 17.7% to 19.8%. Our code is here: https://anonymous.4open.science/r/DyQ-VLA-7F51/.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Ω-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling

    May 27, 2026Xinyu Wang, Mingze Li, Sicheng Lyu +6Diffusion-Based Vision-Language-ActionsGradient Quantization

  2. VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models

    Sep 21, 2026Jiuyi Xu, Qing Jin, Meida Chen +3Post-Training QuantizationClosed Loop

  3. ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

    May 19, 2026Arash Akbari, Arman Akbari, Masih Eskandar +11Post-Training QuantizationFast Inference