cs.ROSep 21, 2026

vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation

Authors: Khanh D. Nguyen, Hoang M. Truong, An T. Le

Organizations: VinRobotics, Vietnam. · Center for AI Research, VinUniversity, Vietnam. · Intelligent Autonomous Systems, TU Darmstadt, Germany.

Abstract

Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query latency and execution horizon to action availability under lagged and time-aligned execution, distinguishing action supply from feedback frequency. Across six policies and four CPUs, vla.simd achieves approximately 1.4×1.4\times median speedup over compiled PyTorch references while preserving fp32 numerical fidelity. We also introduce IMPACT, an ACT-based policy with cached text representations and language-modulated visual features. IMPACT is the only language-conditioned policy in our evaluated set that supplies at least 30 actions/s on the Raspberry Pi 5: after a 90 s thermal soak, it supplies 33.5 actions/s in fp32 and 81.2 with int8. Separate GPU evaluations yield 76.4%76.4\% mean success across four LIBERO suites without robot pretraining; instruction-shuffling tests demonstrate selection among familiar goals. Trials with IMPACT on an SO-101 arm and SmolVLA on a UR10e with a Robotiq gripper demonstrate CPU deployment on two robot embodiments.

Figures & tables

Explore similar work

CardsList
  1. vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

    Jun 6, 2026Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen +5Vision-Language-Action FrameworkFast Inference

  2. Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

    Jul 14, 2026Yuzhou Wu, Yuxin Zheng, Muchun Niu +6Diffusion-Based Vision-Language-ActionsVision-Language-Action Framework

  3. Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies

    Sep 16, 2026Xiatao Sun, Chen Liang, Ziyao Zeng +5DecouplingAction Chunks