Organizations: School of Computer Science, Peking University, Beijing, China. · School of Software Engineering, South China University of Technology, Guangzhou, China. · School of Artificial Intelligence, Beijing Normal University, Beijing, China. · School of Electronics Engineering and Computer Science, Peking University, Beijing China.
Vision-language-action (VLA) models achieve favorable task performance, yet runtime errors in closed-loop execution evolve with alternating updates of actions and observations. Existing VLA quantization methods mainly trade off inference speed and model performance while rarely investigating how quantization alters runtime errors. By comparing error distributions between full-precision and quantized models, we identify distinct evolutionary patterns for the two types of errors: quantization enlarges the variance of translation error distributions and increases their dispersion, whereas the distribution center of rotation errors gradually shifts across execution steps, demonstrating cumulative drift. Motivated by this observation, we rethink the optimal quantization strategy for VLA models and propose \textit{DyQ-VLA}, a runtime-error-aware quantization framework, which dynamically selects activation precision according to execution steps and tracks as well as compensates rotation errors via accumulated quantization residuals. Corresponding operators and runtime adaptation strategies are devised within the framework to enable dynamic-precision execution. Experiments show that \textit{DyQ-VLA} achieves a 1.88 to 1.93 inference speedups while maintaining comparable or higher average task success rates. Moreover, it reduces mean execution steps by 17.7% to 19.8%. Our code is here: https://anonymous.4open.science/r/DyQ-VLA-7F51/.
Figures & tables
Figure 1: Action-Error Correlation and DyQ-VLA Overview. (a) Runtime errors and actions interaction of VLA Models. (b)Runtime-Error-Aware VLA Quantization Framework – DyQ-VLA .
Figure 2: Action–Error Correlation in Closed-Loop VLA Execution. (a) Runtime errors alter observations and subsequent actions, which in turn change runtime errors. (b) Joint distributions of correlation changes in both center and shape across steps. (c) The distribution for translation remains dispersed around a stable center, while rotation stays concentrated and shifts toward larger errors.
Figure 3: Quantization Effects and Online Tracking. (a) Quantization increases translation error variance and response sensitivity. (b) Rotation error distributions and sensitivity change only slightly. (c) Residual and feedback metrics track rotational error and translational sensitivity, respectively.
Figure 4: Framework of DyQ-VLA . Runtime feedback guides dynamic activation quantization for translation error control, while cumulative residuals guide rotation error compensation.
Method
Prec.
RoboTwin 2.0 / π0.5
LIBERO
Clean
Randomized
Steps ↓
Speedup ↑
π0.5
GR00T-N1.5
SR ↑
Trans.
Rot.
SR ↑
Trans.
Rot.
Step
Rollout
Step
Rollout
Mem. ↓
Step
Rollout
Mem. ↓
BF16 ref.
BF16
82.7
32.53
8.77
76.8
39.82
12.27
155.7
1.00
1.00
1.00
1.00
6.71
1.00
1.00
5.45
ActQuant
W4A16
80.9
41.65
12.10
74.2
52.38
16.48
186.4
1.11
0.93
1.10
0.92
2.63
1.09
0.94
4.03
QVLA
W4A16
80.5
42.18
10.92
73.7
51.74
15.28
188.2
1.09
0.90
1.08
0.89
2.70
1.07
0.93
4.10
QuantVLA
W4A8
81.4
40.36
15.24
74.9
50.81
20.31
181.7
1.30
1.11
1.32
1.13
3.50
1.23
1.09
4.27
Table 2: Task Performance and Inference Efficiency on RoboTwin 2.0 and LIBERO.
Figure 5: Real-World Setup and Task Examples. Details can be found in Appendix F .
Method
Prec.
Short-horizon
Multi-stage
SR (%) ↑
Task Spd. ↑
Banana
Cup
Banana
Apple
Orange
Ordered
Plate
Insertion
Cup
Plate
Packing
Placement
BF16 ref.
BF16
9/10
6/10
8/10
8/10
7/10
7/10
75.0
1.00
QuantVLA
W4A8
8/10
6/10
8/10
8/10
6/10
5/10
68.3
1.16
ActQuant
W4A16
8/10
6/10
8/10
8/10
6/10
5/10
68.3
1.08
QVLA
W4A16
8/10
4/10
8/10
8/10
5/10
5/10
63.3
1.04
Table 3: Real-world performance of DyQ-VLA , quantization baselines, and the BF16 reference.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Trajectory Visualizations of Runtime Error Dynamics and Their Behavioral Impact. (a) FP reference, open-loop evaluation, and closed-loop execution trajectories. (b) Identical action offsets at different steps, with trajectories projected onto the yz plane. (c) Offsets of equal magnitude in different directions at the same step.
Figure 9: Action-Error Distributions Across Models and Benchmarks.
Figure 10: Tracking Rotational Runtime Error. (a) Spearman correlation with rotational runtime error. (b) Spearman correlation between signal changes and error changes. (c) NRMSE of normalized changes for error decreases ≥0.5∘ ; the dashed line marks the zero-change baseline. Dots denote trajectories; boxes show medians and interquartile ranges, with 1.5IQR whiskers.
Figure 12: Implementation details of the Proposed DyQ-VLA Framework.
Figure 13: Error Densities.
Figure 14: Real-World Platform And End-Effector Configurations. (a) Tabletop Workspace And Hardware Components. (b) Configurations for Demonstration Collection And Task Execution.
Resource
Setting
Scope
Primary Role
OXE
Real-world
Demonstrations across multiple robot embodiments
Robot learning data
RoboCasa
Simulation
Household manipulation in kitchen environments
Task environments and demonstration data
SimplerEnv
Simulation
Manipulation setups corresponding to real robots
Policy evaluation in simulation
Our setup
Real-world
Six tabletop tasks on a single robotic platform
Task success and completion-time evaluation
Appendix
Table 6: Overview of Robotic Data And Evaluation Settings.
Task Group
Task
Objective
Short-horizon
Banana plate
Place the banana on the beige plate.
Cup insertion
Insert one cup into the other cup.
Banana cup
Insert the banana into the cup.
Apple plate
Place the apple on the plate.
Multi-stage
Orange packing
Place the orange in the case and close the lid.
Ordered placement
First place the apple on the beige plate, then place the banana on the blue plate.
Appendix
Table 7: Real-World Tasks And Objectives.
Figure 15: Representative Real-World Task Sequences. The examples show banana insertion, apple placement, and ordered placement of the apple and banana on their respective plates.
Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive. Prior quantization efforts offer only partial solutions, compressing the LLM backbone while leaving the DiT action head at full precision, or resorting to mixed-precision schemes, driven by the belief that uniformly quantizing the action head is inherently unstable. We challenge this assumption with Omega-QVLA, the first training-free post-training quantization framework that compresses both the language backbone and the entire diffusion action head of a VLA model to a uniform W4A4 precision, eliminating the need for mixed-precision allocation. Omega-QVLA combines a composite SVD-Hadamard rotation that equalizes per-channel weight energy while diffusing residual activation outliers with per-step DiT activation scaling quantization that absorbs dynamic-range drift across denoising steps. On LIBERO, Omega-QVLA compresses Pi 0.5 and GR00T N1.5 to W4A4 with 98.0% and 87.8% task success rates, matching or exceeding their FP16 references of 97.1% and 87.0%, while reducing the static memory footprint by 71.3%. Real-world manipulation experiments further confirm smooth, accurate manipulation where prior methods fail. Code is available at https://github.com/UCMP13753/Omega-QVLA.
Xinyu Wang, Mingze Li, Sicheng Lyu +6
McGill University · Université de Montréal · Mila – Quebec AI Institute +2
Post-training quantization reduces the memory requirements of vision-language-action (VLA) models, but precision selection must account for the interaction between layer scope, numerical format, and calibration. We introduce \textbf{VLAQuantBench}, a controlled evaluation with 409 runs and 94,574 simulation episodes: four models on LIBERO, with X-VLA additionally evaluated on three simulation benchmark families. Under uncalibrated W4A4 round-to-nearest quantization, expanding a π0.5 action-head subset from 126 to 167 layers raises success from 7.0% to 70.5%. Fixed-observation replay confirms a corresponding numerical recovery. Two-episode calibration removes the severe joint failures in the tested subsets, whereas the same smoothing-and-clipping recipe lowers π0 success and does not recover OpenVLA-OFT end-to-end. For OpenVLA-OFT, protecting one 28,672-parameter output projection instead restores near-baseline success: the remaining 441 eligible linear layers retain W3 on LIBERO-Long or eight-bit activations across all four suites. Task-clustered intervals support the large failure and recovery contrasts. These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules. Real-kernel and physical-robot measurements complement the accuracy analysis. Code, configurations, and episode records are publicly available at https://github.com/jiuyixu25/VLAQuantBench.
Jiuyi Xu, Qing Jin, Meida Chen +3
Colorado School of Mines · Independent Researcher · University of Central Florida +2
Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressive, sub-4-bit weight quantization is the natural solution, yet existing post-training quantization (PTQ) methods suffer severe performance degradation in this regime. To address this, we introduce ActQuant, an action-guided mixed-precision PTQ framework that operates in two stages: (1) an inter-tensor bit allocator that assigns each weight matrix a single bit-width based on how much it contributes to predicting the agent's actions; (2) an intra-tensor scale optimizer tunes per-block quantization scales using action-aware curvature, so that dynamic range is concentrated on the weights most influential for control. To deliver the on-device benefits of our aggressive quantization, we further introduce OmniModel.cpp, an agentic conversion pipeline that ports architectures into a native C/C++ runtime with efficient low-bit kernels. We evaluate ActQuant both in simulation and on a real-world 6-DoF UR3 arm, with all models deployed through OmniModel.cpp. On the LIBERO benchmark, ActQuant is the only method that operates at or below 3 bits-per-weight, retaining 95.0% on OpenVLA-OFT and 94.8% on π0.5. Pushed further, ActQuant reaches 2.5 bpw at 90.1% on OpenVLA-OFT, compressing the backbone from 14.3 GB to 2.7 GB (5.3×). On the physical UR3 arm, π0.5 quantized with ActQuant retains the baseline's success rate while reducing the memory footprint by 2.5×.
Arash Akbari, Arman Akbari, Masih Eskandar +11
1Northeastern University · University of Georgia · 3Cisco Systems