cs.ROOct 6, 2026

ActTune: Action-Aware Precision and GPU Operating-Point Adaptation for Energy-Efficient Vision-Language-Action Inference

Authors: Zou Qingyun, Bin Gao, Wenju Zhao, Weng-Fai Wong, Bingsheng He, Tulika Mitra

Organizations: National University of Singapore · Huazhong University of Science and Technology

Abstract

Vision-language-action (VLA) policies repeatedly invoke inference to control robots, making graphics processing unit (GPU) energy a recurring cost of task execution. Reducing energy per inference call, however, may not reduce energy per successful task if numerical errors increase failures or slower inference prolongs execution. We therefore target GPU energy per successful task while preserving task success and keeping the inference-latency increase within 10%. Our approach builds on two observations: quantization sensitivity varies across action classes, model layers, and weights versus activations; and numerical precision changes the workload, shifting favorable GPU operating points. We introduce ActTune, an action-aware framework that connects layer-wise precision allocation with workload-dependent GPU operating-point selection over requested frequency--power-cap pairs. A lightweight decision tree learns its splits and leaf precision configurations directly from configuration action errors, then selects precision before each policy call. The controller forecasts the next workload and applies the selected GPU operating point asynchronously using a lookup table calibrated under a latency budget. A shared resident quantized weight bank enables configuration switching without weight reconstruction or additional policy evaluations. On LIBERO, a benchmark for lifelong robot learning, ActTune improves mean task success by up to 2.3% relative to state of the art. Relative to the original BF16 implementations, it delivers up to 2.02×2.02\times faster inference and, with GPU operating-point adaptation, reduces energy per successful task by up to 76.8%.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Oct 4, 2026cs.AR

Beyond LLM Serving: Characterizing Vision-Language-Action Workloads for Embodied AI System Design

Vision-language-action (VLA) models translate multimodal observations into low-level robot actions. During robot operation, each control period sets an inference deadline, and overruns leave the robot acting on stale observations, reducing task success. Meeting this deadline motivates on-device or nearby edge execution, where a single robot requires batch-1 inference outside the design point of LLM serving systems. Although VLA architectures combine familiar vision-language, autoregressive, and diffusion-style components, their runtime behavior in this batch-1 control setting remains uncharacterized. We characterize four representative VLA models on an edge GPU server and two onboard SoCs, using single-inference profiling and 43,200 closed-loop episodes. Action tensor dimensionality determines whether a stage is memory- or compute-bound, platform balance can shift that bottleneck, and GPU frequency scaling yields a platform-dependent energy-latency sweet spot. In closed-loop operation, overlapping inference with action execution creates an accuracy-speed-energy tradeoff, and no configuration is Pareto-dominant across deployment SLOs. These results guide joint design of VLA model architectures, hardware, and runtime policies.
Sep 21, 2026cs.RO

FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding

Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and π0.5π_{0.5}, W4A4 achieves 1.201.20 to 1.33×1.33\times speedups over floating-point TensorRT on Orin and 1.251.25 to 1.52×1.52\times on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from 80.0%80.0\% with uniform W4A4 to 92.5%92.5\% over 80 trials per configuration, with a measured additional Orin latency of 1 ms.
Jun 6, 2026cs.RO

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Deploying vision--language--action (VLA) models on robots requires adapting model-specific inference pipelines to heterogeneous processors and limited onboard memory. We present vla.cpp, a unified C++ inference runtime for eleven VLA models, with no PyTorch dependency for model execution. The runtime shares model loading, tensor execution, and serving while retaining architecture-specific attention, conditioning, and action heads. Iterative policies reuse observation-dependent computation across solver steps, while regression policies predict actions directly. We evaluate task success on LIBERO-Object and profile supported configurations on NVIDIA, Apple, and Intel hardware. BitVLA completes 200/200 LIBERO-Object episodes on an 8,GB Jetson Orin Nano. A ternary tensor-core kernel accelerates its client inference by 4.04.0--4.6×4.6\times over the CUDA-core baseline on RTX 3060 and AGX Orin. A SmolVLA case study links positional-index precision to gripper commands and task success, showing why fixed-input numerical checks should accompany rollout evaluation. Deployments on UR10e and ALOHA demonstrate physical robot integration; delay and execution-horizon studies characterize synchronous chunked control. The results demonstrate a common deployment path across VLA architectures and hardware, with numerical validation and control settings guiding deployment alongside inference efficiency.