Vision-language-action (VLA) models translate multimodal observations into low-level robot actions. During robot operation, each control period sets an inference deadline, and overruns leave the robot acting on stale observations, reducing task success. Meeting this deadline motivates on-device or nearby edge execution, where a single robot requires batch-1 inference outside the design point of LLM serving systems. Although VLA architectures combine familiar vision-language, autoregressive, and diffusion-style components, their runtime behavior in this batch-1 control setting remains uncharacterized. We characterize four representative VLA models on an edge GPU server and two onboard SoCs, using single-inference profiling and 43,200 closed-loop episodes. Action tensor dimensionality determines whether a stage is memory- or compute-bound, platform balance can shift that bottleneck, and GPU frequency scaling yields a platform-dependent energy-latency sweet spot. In closed-loop operation, overlapping inference with action execution creates an accuracy-speed-energy tradeoff, and no configuration is Pareto-dominant across deployment SLOs. These results guide joint design of VLA model architectures, hardware, and runtime policies.
Figures & tables
Latency
NCU GPU time (%)
Stage
(ms)
Memory- bound
Compute- bound
Vision encoder
16.03
23.9
76.1
LLM prefill
43.50
13.6
86.4
Cached decode (6 passes)
140.90
100.0
0.0
Other + preprocessing
9.13
–
–
E2E
209.56
–
–
Table 1 . Latency decomposition and NCU roofline classification of autoregressive OpenVLA on RTX 4090 in eager mode.
Figure 1 . Architectures and profiling decomposition of the four action-chunk models evaluated in closed loop. Block diagrams compare parallel decoding, two iterative-denoising architectures, and a video-based policy, showing the backbone and action-generation stages used for profiling.
Figure 2 . Synchronous and asynchronous inference. In synchronous mode, the robot idles during inference (top). Asynchronous overlap initiates the next inference before the current chunk completes, with the overlap percentage governing when the next request is triggered (middle and bottom). Timelines compare synchronous execution with asynchronous schedules that trigger the next inference after different fractions of the current action chunk have been consumed.
Model
Backbone
Action Head
Config 1
GR00T N1.6
Cosmos-Reason VLM (Eagle ViT + LLM)
DiT (Flow matching)
4/3B/16
π0.5
PaliGemma VLM (SigLIP+Gemma 2B)
Gemma 300M (Flow matching)
10/3B/10
OpenVLA-OFT
Prismatic VLM (SigLIP+DINOv2+ Llama 2)
MLP (L1 regression)
N/A/7B/8
Cosmos-Policy
Wan2.1 VAE
DiT (Diffusion) 2
5/2B/16
Table 2 . Architecture overview of evaluated models.
Platform
Tier
Memory
Power
SW Stack
Jetson AGX Orin
Power-efficient SoC
64 GB LPDDR5
15–60 W
JetPack 6.2
Jetson Thor
High-performance SoC
128 GB LPDDR5X
70–130 W
JetPack 7.0
RTX 4090
Edge server
24 GB GDDR6X
450 W
CUDA 13.1
Table 3 . Hardware platforms and software stacks used in evaluation.
Analysis
Metric
Tool
Sec.
Latency decomp.
Wall-clock per stage (ms)
perf_counter + cuda.sync
4.1
Roofline profiling
GPU kernel time, FLOPs, DRAM bytes
NVIDIA NCU (v2025.2.0)
4.2
DVFS sensitivity
E2E latency per freq. point
perf_counter + sysfs
4.3
Closed-loop
Success rate, measured energy
MuJoCo + NVML / sysfs
5
Table 4 . Measurement methodology summary. Each analysis targets a different level of the system stack; detailed configurations are deferred to the corresponding sections.
Model
Backbone
Action Head
GR00T N1.6
Default
Max-autotune
π0.5
Max-autotune
Max-autotune
OpenVLA-OFT
Max-autotune *
N/A
Cosmos-Policy
Eager
Eager
* Includes prewarm for variable-length task prompts.
Table 5 . Compilation configuration used for DVFS analysis and closed-loop evaluation. Compilation mode refers to torch.compile mode; “eager” denotes no compilation. Stage-level decomposition and roofline profiling use eager mode throughout.
Figure 3 . End-to-end (E2E) latency breakdown in eager mode. Stacked bars decompose eager-mode inference latency into model stages for the four action-chunk models.
Figure 4 . Per-stage compilation speedup over eager execution by model. Grouped bars compare compilation speedup for the backbone and action-generation stages of each evaluated model.
Figure 5 . Roofline classification by stage on RTX 4090. Mem = memory-bound fraction of GPU time; Comp = compute-bound fraction. Stacked bars divide each model stage's GPU time into memory-bound and compute-bound fractions.
Figure 6 . Operator breakdown by stage and model on the RTX 4090. Stacked bars show the fraction of GPU time spent in GEMM and other operator classes for each model stage.
Figure 7 . GEMM size and arithmetic intensity on the RTX 4090 roofline. A roofline scatter plot places GEMM kernels by operation count and arithmetic intensity, grouped by model and pipeline stage.
Figure 8 . Per-stage GPU/EMC sensitivity ratio R (Eq. 3 ) on Orin and Thor. Points are colored by the roofline-predicted bottleneck (memory- or compute-bound) from RTX 4090 eager-mode profiling. Points compare per-stage sensitivity to GPU and memory frequency on Orin and Thor, colored by the RTX 4090 roofline classification.
Figure 9 . Energy–latency tradeoffs under GPU and EMC frequency sweeps on Jetson AGX Orin and Jetson Thor. Energy-latency curves compare GPU-frequency and memory-frequency sweeps on the two edge SoCs.
Figure 10 . Transition-free latency–energy tradeoffs under the three-rule phase-aware policy on Orin and Thor, normalized to the default platform DVFS configuration (lower is better). Open stars denote energy changes within measurement resolution, and candlesticks show energy variability. Eight panels plot normalized latency against normalized energy for phase-aware GPU-frequency assignments on Orin and Thor across four models.
Transition latency
Share of E2E latency
Platform
Per transition
Per inference
GR00T N1.6
π0.5
Cosmos- Policy
Orin
3.5 ms
7.0 ms
2.5%
1.2%
0.3%
Thor
8.6 ms
18.0 ms
13.6%
7.8%
1.7%
Table 6 . Measured GPU-frequency transition latency and its share of baseline E2E latency for two transitions per inference. OpenVLA-OFT has no intra-inference transition and is omitted.
Control Period (ms)
Model
Platform
RTTmax (ms)
10%
25%
50%
GR00T N1.6
RTX 4090
142.8
107.1
53.5
26.8
Thor
126.7
95.1
47.5
23.8
Orin
206.8
155.1
77.6
38.8
π0.5
RTX 4090
136.1
204.1
68.0
40.8
Thor
180.6
270.8
90.3
54.2
Table 7 . Worst-case round-trip time ( RTTmax ) and control periods per overlap level. RTTmax averages each episode’s maximum RTT under synchronous operation and covers observation transfer, server-side pre- and post-processing, and compiled inference. Control periods are derived via Eq. 5 .
Figure 11 . System architecture of the closed-loop benchmark. The client node runs a MuJoCo simulator and consumes actions from a FIFO queue at fixed intervals Δt , while asynchronously prefetching new action chunks from multi-platform model servers over WebSocket. A client-server diagram shows a MuJoCo client consuming actions from a FIFO queue while asynchronously requesting action chunks from model servers.
Figure 12 . Accuracy–speed–energy tradeoff across all model–platform combinations, split into our two task groups: Long (LIBERO-10) and Short (LIBERO-Goal/Spatial/Object). Each subplot shows success rate (magenta, left axis), average episode time (blue, right axis), and average measured energy (tan bars, annotated in J) as a function of overlap percentage. A grid compares success rate, episode time, and measured energy across overlap levels for every model and platform, with separate curves for Long and Short task groups.
Figure 13 . Task-equal successful-episode energy versus action throughput, per task group and platform, across overlap levels (10%–50%). Each trajectory connects one model’s overlap settings; lower right (higher throughput at lower energy) is better. Panels plot task-equal energy per successful episode against action throughput, with trajectories connecting each model's overlap settings on each platform and task group.
Figure 14 . Deployment operating points under SLO constraints. Shading marks values below each metric’s lower-quartile (p25) budget; callouts mark the highest-SR feasible Long and Short points. Filled markers: Long (LIBERO-10); open markers: Short (LIBERO-Goal/Spatial/Object). Four scatter plots show success rate against episode time, edge power, measured energy, and control-cycle latency, with shaded feasible regions and selected Long and Short operating points.