cs.ROOct 8, 2026

REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models

Authors: Houlong Xiong, Zhenqi Qiu, Zechen Wang, Suohang Zhang, Yiyu Ren, Wanting Xu, Hongfei Niu, Chengyang He, +3 more

Organizations: Shanghai Jiao Tong University · PrimeBot · ShanghaiTech University · Zhejiang University · National University of Singapore

Abstract

Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReactVLA: Fast and Lightweight Reactive Robot Manipulation via Improved Mean Flow Action Generation

    Jun 12, 2026Yanzhao Guo, Wenkai Chen, Jianwei ZhangLanguage-Conditioned Robot ManipulationEfficient VLM Inference

  2. RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models

    Sep 28, 2026Yuhan Chen, Ke Yu, Pengfei Liu +3Efficient VLA Model InferenceVision-Language-Action Models