cs.ROOct 1, 2026

ChunkVLA-AM: Parallel Action Chunking for Vision-Language-Action Robot Control in Additive Manufacturing

Authors: Zhugang Liu, Kaichuang Zhang, Jinman Zhang, Pu Sun, Martha Asare, Jose Hernandez, Maxim Ermolinsky, Efren Saenz, +2 more

Organizations: Department of Computer Science, The University of Texas Rio Grande Valley · Department of Electrical Engineering, University of South Florida · Department of Computer Science, San Diego State University · Department of Electrical and Computer Engineering, The University of Texas Rio Grande Valley

Abstract

Vision-language-action (VLA) models unify visual perception, language understanding, and action generation, offering new opportunities for automation in additive manufacturing (AM). However, deployment in AM remains challenging because adapting these models to unseen robot embodiments is costly, and performance can degrade under environment changes. In this work, we present a framework for deploying OpenVLA-OFT on a FAIRINO FR3 robot in a fixed AM workcell. A data pipeline converts monocular real-world demonstrations into OpenVLA-compatible TFDS/RLDS datasets to support adaptation to the FR3 embodiment. At runtime, each inference request predicts an eight-step chunk of 7-D actions. The FR3 executes each chunk open loop before capturing a new observation, providing closed-loop feedback between chunks. The system uses a cloud-edge architecture in which the FR3 client streams observations to a remote inference server through a FastAPI interface. In 42 physical A-to-B object-transfer trials, evenly split between red and blue targets, the system succeeded in 39 (92.9%). All three failures occurred during final placement, when insufficient release-height control caused the object to topple. An illumination sweep identified a low-error luminance range of 85-125 on a 0-255 scale, with the lowest mean spatial error at 95.

Figures & tables

Explore similar work

CardsList
  1. When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models

    Oct 5, 2026Seonghoon Yu, Dongwon Kim, HyungRok Jung +4

  2. VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

    Jul 7, 2025Juyi Lin, Amir Taherin, Arash Akbari +11Diffusion-Based Vision-Language-ActionsRobotic Manipulation

  3. On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation

    Jul 11, 2026Finn Ferchau, Daniel Pommer, Cristian AxenieDiffusion-Based Vision-Language-ActionsFlow-Matching Vision-Language-Action