cs.AROct 4, 2026

Beyond LLM Serving: Characterizing Vision-Language-Action Workloads for Embodied AI System Design

Authors: Seonghun Jung, Sieun Moon, Jiyoung Jeong, Jimin Lee, Jaehyuk Huh

Organizations: KAIST Daejeon, Republic of Korea

Abstract

Vision-language-action (VLA) models translate multimodal observations into low-level robot actions. During robot operation, each control period sets an inference deadline, and overruns leave the robot acting on stale observations, reducing task success. Meeting this deadline motivates on-device or nearby edge execution, where a single robot requires batch-1 inference outside the design point of LLM serving systems. Although VLA architectures combine familiar vision-language, autoregressive, and diffusion-style components, their runtime behavior in this batch-1 control setting remains uncharacterized. We characterize four representative VLA models on an edge GPU server and two onboard SoCs, using single-inference profiling and 43,200 closed-loop episodes. Action tensor dimensionality determines whether a stage is memory- or compute-bound, platform balance can shift that bottleneck, and GPU frequency scaling yields a platform-dependent energy-latency sweet spot. In closed-loop operation, overlapping inference with action execution creates an accuracy-speed-energy tradeoff, and no configuration is Pareto-dominant across deployment SLOs. These results guide joint design of VLA model architectures, hardware, and runtime policies.

Figures & tables

Explore similar work

CardsList
  1. vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

    Jun 6, 2026Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen +5Vision-Language-Action FrameworkFast Inference

  2. Efficient Vision-Language-Action Management and Serving for Robot Factories

    Sep 14, 2026Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri +1Diffusion-Based Vision-Language-Actions