Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query latency and execution horizon to action availability under lagged and time-aligned execution, distinguishing action supply from feedback frequency. Across six policies and four CPUs, vla.simd achieves approximately 1.4× median speedup over compiled PyTorch references while preserving fp32 numerical fidelity. We also introduce IMPACT, an ACT-based policy with cached text representations and language-modulated visual features. IMPACT is the only language-conditioned policy in our evaluated set that supplies at least 30 actions/s on the Raspberry Pi 5: after a 90 s thermal soak, it supplies 33.5 actions/s in fp32 and 81.2 with int8. Separate GPU evaluations yield 76.4% mean success across four LIBERO suites without robot pretraining; instruction-shuffling tests demonstrate selection among familiar goals. Trials with IMPACT on an SO-101 arm and SmolVLA on a UR10e with a Robotiq gripper demonstrate CPU deployment on two robot embodiments.
Figures & tables
Device
CPU core
ISA
Cores
Kernel
Thr.
Apple M4
Apple
NEON
4P+6E
Accelerate
8
Intel i9-14900HX
Raptor Lake
AVX2
8P+16E
SMK 6×16
16
AMD Ryzen 5 5500
Zen 3
AVX2
6C/12T
SMK 6×16
12
Raspberry Pi 5
Cortex-A76
NEON
4C
SMK 4×16
4
TABLE I: CPU platforms and benchmark settings. P/E: performance/efficiency cores; C/T: cores/hardware threads.
Policy
Params
Vision / language
Action model
H
ACT [ 5 ]
34 M
ResNet-18 / -
CVAE transformer
100
DP [ 6 ]
278 M
ResNet-18 / -
U-Net, 100/10 steps
32
Octo-Small [ 3 ]
27 M+T5
conv stem / T5-base
diffusion, 20 steps
4
TurboVLA [ 11 ]
0.2 B
DINOv3 / BERT
ACT-style decoder
12
SmolVLA [ 4 ]
450 M
SmolVLM2
flow matching
50
IMPACT (ours)
60 M+T5
ResNet-18 / T5-small
CVAE transformer
50
TABLE II: Evaluated policies and available action windows H . “+T5” denotes a separate frozen text encoder.
Raspberry Pi 5
Nominal
90 s soak
Policy
H
M4
i9
Ryzen
feff
Tq (s)
fp32
int8
ACT
100
1150
892
627
109
0.92
89.9
242
DP, 100 steps
32
7.2
7.2
3.6
0.8
40.1
-
-
DP, 10 steps
32
58
55
29.3
6.3
5.08
-
-
Octo-Small
4
76
48
48
7.8
0.51
5.8
11.3
TABLE III: Action-supply rate feff (actions/s). Gray : fails C1; italic : meets C1 only; upright black: meets C2, at fc=30 Hz. Nominal: fp32; last two columns: after a 90 s Pi 5 soak; dashes: unmeasured.
Setting
A / B
M4
i9
Ryzen
Pi 5
Attention kernel
library / custom
A + 9.9
=
=
=
bf16 expansion
off / on
=
=
=
=
Min. panel rows
16 / 4
=
=
B − 4.2
=
Convolution
tiled / flat
=
A + 156
A + 260
A + 16.9
Block rows
auto / 64
=
A + 8.3
A + 5.0
=
Thread schedule
static / dyn.
=
=
=
=
TABLE IV: Effects of individual runtime settings on ACT. A/B: consistently faster setting; =: no consistent difference. Values are median 100(TB/TA−1) (%); positive favors A. Bottom: Tfp32/TW8A8 ; above 1 is faster; n/a: unsupported; –: unmeasured.
Policy
Spatial
Object
Goal
Long
Avg
No robot pretraining
Diffusion Policy
78.3
92.5
68.3
50.5
72.4
TurboVLA
99.2
99.8
97.4
94.2
97.7
SmolVLA
90.0
96.0
92.0
71.0
87.3
IMPACT, final
83.5
83.5
84.5
54.0
76.4
IMPACT, selected
83.5
83.5
89.0
58.5
78.6
TABLE V: LIBERO success rates (%); protocols differ across published baselines. IMPACT: final checkpoint versus selection on the same evaluation episodes.
Fig. 5: Real-robot tasks on two embodiments. Top: three instructions in a shared SO-101 [ 38 ] scene vary the destination or grasp target. Middle: the SO-101 opens a drawer, places the tape inside, and closes it. Bottom: a UR10e with a Robotiq gripper picks up the cup and places it in the box. The middle and bottom rows show successive stages from left to right.
Deploying vision--language--action (VLA) models on robots requires adapting model-specific inference pipelines to heterogeneous processors and limited onboard memory. We present vla.cpp, a unified C++ inference runtime for eleven VLA models, with no PyTorch dependency for model execution. The runtime shares model loading, tensor execution, and serving while retaining architecture-specific attention, conditioning, and action heads. Iterative policies reuse observation-dependent computation across solver steps, while regression policies predict actions directly. We evaluate task success on LIBERO-Object and profile supported configurations on NVIDIA, Apple, and Intel hardware. BitVLA completes 200/200 LIBERO-Object episodes on an 8,GB Jetson Orin Nano. A ternary tensor-core kernel accelerates its client inference by 4.0--4.6× over the CUDA-core baseline on RTX 3060 and AGX Orin. A SmolVLA case study links positional-index precision to gripper commands and task success, showing why fixed-input numerical checks should accompany rollout evaluation. Deployments on UR10e and ALOHA demonstrate physical robot integration; delay and execution-horizon studies characterize synchronous chunked control. The results demonstrate a common deployment path across VLA architectures and hardware, with numerical validation and control settings guiding deployment alongside inference efficiency.
Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen +5
VinRobotics, Vietnam · Max Planck Research School for Intelligent Systems, Germany · University of Stuttgart, Germany +3
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2
Vision-Language-Action (VLA) policies commonly run Vision-Language Model (VLM) backbones with billions of parameters at every policy inference, which costs latency and energy. We revisit a decoupled alternative for multi-task manipulation: separate vision and language encoders whose representations condition a compact action head. We run a standardized comparison that varies the vision encoder, the language encoder, and the action head while holding the demonstrations, the training-step budget, the tasks, the evaluation protocol, and the measurement platform fixed, against seven VLA baselines. The resulting Decoupled Embodiment Model (DEM) combines a fine-tuned DINOv3 vision encoder, a frozen NeoBERT language encoder, and a MeanFlow head that generates an action chunk in one forward pass. On 18 RoboCasa tasks evaluated with held-out instruction paraphrases and randomized scenes, DEM reaches 55.6% mean success against 56.9% for GR00T N1.7 and 54.6% for π0.5, and on three real-robot tasks it reaches 66.0% against 68.0% for GR00T N1.7. On the same workstation, DEM needs 6.1,ms per policy forward pass, a maximum throughput of 162.7 policy calls per second, and draws an estimated 2.07,J of GPU energy per call, eight to seventeen times the throughput and six to fifteen times less energy than these VLM-backbone policies. Within this trained-task regime, DEM sits on the observed success--latency--energy frontier and provides a strong, efficient baseline for language-conditioned robot skills.
Xiatao Sun, Chen Liang, Ziyao Zeng +5
Department of Computer Science, Yale University, New Haven, CT, USA. · Peking University, Beijing, China. · Digients, Singapore.