Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query latency and execution horizon to action availability under lagged and time-aligned execution, distinguishing action supply from feedback frequency. Across six policies and four CPUs, vla.simd achieves approximately 1.4× median speedup over compiled PyTorch references while preserving fp32 numerical fidelity. We also introduce IMPACT, an ACT-based policy with cached text representations and language-modulated visual features. IMPACT is the only language-conditioned policy in our evaluated set that supplies at least 30 actions/s on the Raspberry Pi 5: after a 90 s thermal soak, it supplies 33.5 actions/s in fp32 and 81.2 with int8. Separate GPU evaluations yield 76.4% mean success across four LIBERO suites without robot pretraining; instruction-shuffling tests demonstrate selection among familiar goals. Trials with IMPACT on an SO-101 arm and SmolVLA on a UR10e with a Robotiq gripper demonstrate CPU deployment on two robot embodiments.
Figures & tables
Device
CPU core
ISA
Cores
Kernel
Thr.
Apple M4
Apple
NEON
4P+6E
Accelerate
8
Intel i9-14900HX
Raptor Lake
AVX2
8P+16E
SMK 6×16
16
AMD Ryzen 5 5500
Zen 3
AVX2
6C/12T
SMK 6×16
12
Raspberry Pi 5
Cortex-A76
NEON
4C
SMK 4×16
4
TABLE I: CPU platforms and benchmark settings. P/E: performance/efficiency cores; C/T: cores/hardware threads.
Policy
Params
Vision / language
Action model
H
ACT [ 5 ]
34 M
ResNet-18 / -
CVAE transformer
100
DP [ 6 ]
278 M
ResNet-18 / -
U-Net, 100/10 steps
32
Octo-Small [ 3 ]
27 M+T5
conv stem / T5-base
diffusion, 20 steps
4
TurboVLA [ 11 ]
0.2 B
DINOv3 / BERT
ACT-style decoder
12
SmolVLA [ 4 ]
450 M
SmolVLM2
flow matching
50
IMPACT (ours)
60 M+T5
ResNet-18 / T5-small
CVAE transformer
50
TABLE II: Evaluated policies and available action windows H . “+T5” denotes a separate frozen text encoder.
Raspberry Pi 5
Nominal
90 s soak
Policy
H
M4
i9
Ryzen
feff
Tq (s)
fp32
int8
ACT
100
1150
892
627
109
0.92
89.9
242
DP, 100 steps
32
7.2
7.2
3.6
0.8
40.1
-
-
DP, 10 steps
32
58
55
29.3
6.3
5.08
-
-
Octo-Small
4
76
48
48
7.8
0.51
5.8
11.3
TABLE III: Action-supply rate feff (actions/s). Gray : fails C1; italic : meets C1 only; upright black: meets C2, at fc=30 Hz. Nominal: fp32; last two columns: after a 90 s Pi 5 soak; dashes: unmeasured.
Setting
A / B
M4
i9
Ryzen
Pi 5
Attention kernel
library / custom
A + 9.9
=
=
=
bf16 expansion
off / on
=
=
=
=
Min. panel rows
16 / 4
=
=
B − 4.2
=
Convolution
tiled / flat
=
A + 156
A + 260
A + 16.9
Block rows
auto / 64
=
A + 8.3
A + 5.0
=
Thread schedule
static / dyn.
=
=
=
=
TABLE IV: Effects of individual runtime settings on ACT. A/B: consistently faster setting; =: no consistent difference. Values are median 100(TB/TA−1) (%); positive favors A. Bottom: Tfp32/TW8A8 ; above 1 is faster; n/a: unsupported; –: unmeasured.
Policy
Spatial
Object
Goal
Long
Avg
No robot pretraining
Diffusion Policy
78.3
92.5
68.3
50.5
72.4
TurboVLA
99.2
99.8
97.4
94.2
97.7
SmolVLA
90.0
96.0
92.0
71.0
87.3
IMPACT, final
83.5
83.5
84.5
54.0
76.4
IMPACT, selected
83.5
83.5
89.0
58.5
78.6
TABLE V: LIBERO success rates (%); protocols differ across published baselines. IMPACT: final checkpoint versus selection on the same evaluation episodes.
Fig. 5: Real-robot tasks on two embodiments. Top: three instructions in a shared SO-101 [ 38 ] scene vary the destination or grasp target. Middle: the SO-101 opens a drawer, places the tape inside, and closes it. Bottom: a UR10e with a Robotiq gripper picks up the cup and places it in the box. The middle and bottom rows show successive stages from left to right.