Vision-language-action (VLA) policies repeatedly invoke inference to control robots, making graphics processing unit (GPU) energy a recurring cost of task execution. Reducing energy per inference call, however, may not reduce energy per successful task if numerical errors increase failures or slower inference prolongs execution. We therefore target GPU energy per successful task while preserving task success and keeping the inference-latency increase within 10%. Our approach builds on two observations: quantization sensitivity varies across action classes, model layers, and weights versus activations; and numerical precision changes the workload, shifting favorable GPU operating points. We introduce ActTune, an action-aware framework that connects layer-wise precision allocation with workload-dependent GPU operating-point selection over requested frequency--power-cap pairs. A lightweight decision tree learns its splits and leaf precision configurations directly from configuration action errors, then selects precision before each policy call. The controller forecasts the next workload and applies the selected GPU operating point asynchronously using a lookup table calibrated under a latency budget. A shared resident quantized weight bank enables configuration switching without weight reconstruction or additional policy evaluations. On LIBERO, a benchmark for lifelong robot learning, ActTune improves mean task success by up to 2.3% relative to state of the art. Relative to the original BF16 implementations, it delivers up to 2.02× faster inference and, with GPU operating-point adaptation, reduces energy per successful task by up to 76.8%.
Figures & tables
Figure 1: ActTune intuition: approach tolerates errors, alignment requires accuracy, and transport precision depends on path and grip stability.
Figure 2: Action grouping and precision sensitivity. PC1 and PC2 denote the first and second principal components. W n A m denotes n -bit weights and m -bit activations. Crosses mark cluster centroids; success rates average all four suites, with all action and layer precision batches complete.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Budget
Candidates
Retained
Retained pairs
OpenVLA-OFT
W4A8
3
3
W2A8, W4A8, W8A8
OpenVLA-OFT
W4A4
9
7
W2A2, W2A4, W4A2, W4A4, W4A8, W8A4, W8A8
π0.5
W4A8
3
3
W2A8, W4A8, W8A8
π0.5
W4A4
9
6
W2A2, W2A4, W4A2, W4A4, W4A8, W8A4
Appendix
Table 1: Candidate and retained local weight/activation pairs in the frozen evaluated banks. Retained counts are distinct bit-width pairs, not action classes or whole-model configurations.
Suite
Calibration trajectories
Calibration frames
Diagnostic trajectories
Diagnostic frames
Spatial
135
2,139
10
162
Object
121
2,341
10
185
Goal
138
2,305
10
174
Long
118
4,211
10
361
Total
512
10,996
40
882
Appendix
Table 2: Shared training calibration and held-out diagnostic data.
Model
Budget
Suite
Switchable layers
Precision options
OFT
W4A8
Spatial, Object, Long
V.1.attn.qkv V.22.mlp.fc1
W2A8 / W4A8 W2A8 / W4A8
OFT
W4A8
Goal
V.19.attn.proj V.22.attn.proj
W2A8 / W4A8 W2A8 / W4A8
OFT
W4A4
Spatial
V.1.attn.qkv V.23.mlp.fc1
W2A2 / W4A2 W2A2 / W4A2
OFT
W4A4
Object
V.17.attn.proj V.23.mlp.fc1
W2A2 / W4A2 W2A4 / W4A4
OFT
W4A4
Goal
V.17.attn.proj V.22.mlp.fc1
W2A2 / W4A2 W2A2 / W4A2
OFT
W4A4
Long
V.22.attn.proj V.23.mlp.fc1
W2A4 / W4A4 W2A2 / W4A2
Appendix
Table 3: Switchable layers in the evaluated banks. OFT denotes OpenVLA-OFT; V, L, and E denote its fused vision encoder, the π0.5 language model, and the π0.5 action expert, respectively. Block indices start at zero. Precision options correspond to the listed layers in order.
Figure 3: OFT W4A8 leaf-count ablation for one tree shared across four LIBERO suites. Bars show validation action-loss reduction relative to one leaf; the line shows CPU routing latency. Three leaves give the lowest observed validation loss in this setting.
Figure 4: Shared-tree leaf-count ablations for the remaining model–precision settings. Bars: validation action-loss reduction; lines: CPU routing latency. Three leaves are highlighted for comparison.
Model
Setting
Spatial
Object
Goal
Long
Mean
Speedup
OpenVLA-OFT
BF16
149.21
150.69
148.71
151.52
150.03
1.00×
OpenVLA-OFT
W4A8
99.76
99.65
99.47
101.31
100.05
1.50×
OpenVLA-OFT
W4A4
98.50
98.08
98.16
97.87
98.15
1.53×
π0.5
BF16
239.00
243.90
243.76
247.59
243.56
1.00×
π0.5
W4A8
119.72
120.38
120.22
121.23
120.39
2.02×
π0.5
W4A4
133.97
134.15
135.47
135.69
134.82
1.81×
Appendix
Table 4: Per-suite mean inference latency (ms) and speedup relative to the original BF16 runtime.
Model
Setting
BF16
Earlier
Optimized
Speedup vs. earlier
Speedup vs. BF16
OpenVLA-OFT
W4A8
150.03
134.33
100.05
1.34×
1.50×
OpenVLA-OFT
W4A4
150.03
100.32
98.15
1.02×
1.53×
π0.5
W4A8
243.56
373.39
120.39
3.10×
2.02×
π0.5
W4A4
243.56
398.94
134.82
2.96×
1.81×
Appendix
Table 5: Full-policy inference latency before and after runtime optimization (ms), with speedups relative to the earlier quantized runtime and original BF16 implementation.
Model
Precision
Quantization
Kernel + runtime
DVFS
Total (%)
OpenVLA-OFT
W4A8
49.82
9.71
10.68
70.21
OpenVLA-OFT
W4A4
59.27
0.61
8.14
68.02
π0.5
W4A8
28.09
42.82
5.92
76.83
π0.5
W4A4
−27.73
73.88
10.77
56.92
Appendix
Table 6: Energy-saving contributions from sequential ablation measurements, with every component normalized by the corresponding measured BF16 energy per success. Component values are percentage points; total savings are percentages. The displayed components sum to the total saving in each row.
Intervention
C0
C1
C2
BF16
95.0
95.0
95.0
W2A2
62.5
35.0
91.9
W2A4
67.5
42.5
95.0
W2A8
97.5
97.5
94.4
W4A2
75.0
62.5
95.0
W4A4
70.0
42.5
92.5
Appendix
Table 7: Final success after one action-block precision intervention (%). Entries average the four suites; the matched BF16 mean is 95.0%.
Vision-language-action (VLA) models translate multimodal observations into low-level robot actions. During robot operation, each control period sets an inference deadline, and overruns leave the robot acting on stale observations, reducing task success. Meeting this deadline motivates on-device or nearby edge execution, where a single robot requires batch-1 inference outside the design point of LLM serving systems. Although VLA architectures combine familiar vision-language, autoregressive, and diffusion-style components, their runtime behavior in this batch-1 control setting remains uncharacterized. We characterize four representative VLA models on an edge GPU server and two onboard SoCs, using single-inference profiling and 43,200 closed-loop episodes. Action tensor dimensionality determines whether a stage is memory- or compute-bound, platform balance can shift that bottleneck, and GPU frequency scaling yields a platform-dependent energy-latency sweet spot. In closed-loop operation, overlapping inference with action execution creates an accuracy-speed-energy tradeoff, and no configuration is Pareto-dominant across deployment SLOs. These results guide joint design of VLA model architectures, hardware, and runtime policies.
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and π0.5, W4A4 achieves 1.20 to 1.33× speedups over floating-point TensorRT on Orin and 1.25 to 1.52× on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from 80.0% with uniform W4A4 to 92.5% over 80 trials per configuration, with a measured additional Orin latency of 1 ms.
Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen +5
VinRobotics, Vietnam · AICV Lab, EECS Department, University of Arkansas · School of Advanced Manufacturing and Robotics, Peking University, China +2
Deploying vision--language--action (VLA) models on robots requires adapting model-specific inference pipelines to heterogeneous processors and limited onboard memory. We present vla.cpp, a unified C++ inference runtime for eleven VLA models, with no PyTorch dependency for model execution. The runtime shares model loading, tensor execution, and serving while retaining architecture-specific attention, conditioning, and action heads. Iterative policies reuse observation-dependent computation across solver steps, while regression policies predict actions directly. We evaluate task success on LIBERO-Object and profile supported configurations on NVIDIA, Apple, and Intel hardware. BitVLA completes 200/200 LIBERO-Object episodes on an 8,GB Jetson Orin Nano. A ternary tensor-core kernel accelerates its client inference by 4.0--4.6× over the CUDA-core baseline on RTX 3060 and AGX Orin. A SmolVLA case study links positional-index precision to gripper commands and task success, showing why fixed-input numerical checks should accompany rollout evaluation. Deployments on UR10e and ALOHA demonstrate physical robot integration; delay and execution-horizon studies characterize synchronous chunked control. The results demonstrate a common deployment path across VLA architectures and hardware, with numerical validation and control settings guiding deployment alongside inference efficiency.
Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen +5
VinRobotics, Vietnam · Max Planck Research School for Intelligent Systems, Germany · University of Stuttgart, Germany +3