Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.
Figures & tables
Figure 1: Overview of REACT. REACT reformulates flow-based action generation into a receding-horizon rolling denoising process. At each control step, the active horizon is refined conditioned on the latest observation embedding; the oldest refined chunk is executed, while the remaining buffer is shifted forward and replenished with fresh noise. The colored borders track the same action chunk as it is progressively refined across consecutive observations before execution. The right panels summarize throughput and overall performance, showing that REACT reaches the high-input, high-output regime while improving success, reactivity, smoothness, and efficiency.
Figure 2: Real-world general manipulation tasks. ARX X5 (top) and Franka R3 (bottom) perform two shared tasks, Bowl Stacking by Size and Bottle Cap Unscrewing, plus robot-specific hanging tasks: Cable Hanging and Keyring Hanging, respectively.
π0.5 H50E50
π0.5 H50E10
π0.5 H10E10
Async+RTC
REACT w/o DD
REACT
Task
Clean
Rand
Clean
Rand
Clean
Rand
Clean
Rand
Clean
Rand
Clean
Rand
beat_block_hammer
17
1
8
0
0
0
13
4
15
2
14
4
click_bell
61
6
71
18
35
0
64
8
89
32
61
22
move_playingcard_away
59
26
71
13
7
1
64
11
71
25
66
22
pick_diverse_bottles
20
3
9
0
0
1
10
6
27
9
33
6
press_stapler
58
23
22
3
9
4
37
13
67
42
63
39
Table 1: Simulation results on RoboTwin 2.0 benchmark. Success rate (%) under clean and randomized settings. Best in bold , second best underlined .
Robot
Task
π0.5 H50E50
π0.5 H50E10
π0.5 H10E10
Async+RTC
REACT w/o DD
REACT
ARX X5
Bowl Stacking by Size
73
66
70
36
86
76
Bottle Cap Unscrewing
36
0
0
0
56
47
Cable Hanging
56
66
56
3
63
80
ARX X5 Avg.
55.0
44.0
42.0
13.0
68.3
67.7
Franka
Bowl Stacking by Size
60
50
0
47
66
63
Bottle Cap Unscrewing
57
0
0
3
63
70
Table 2: Real-world results on general manipulation tasks. Success rate (%) across ARX X5 and Franka platforms. Best in bold , second best underlined .
Figure 3: Trajectory smoothness comparison. Mean jerk norm averaged over successful episodes in simulation, ARX X5, and Franka settings. Demos (GT) applies the same metric to the imitation-learning demonstrations as a reference. Full per-task values are reported in Appendix D .
Figure 4: Specialized dynamic closed-loop control tasks. (a) Pour Rice requires the robot to pour rice into a vessel and stop at the target state, testing whether the policy can react to continuously changing flow while maintaining smooth motion. (b) Reaction Game requires the robot to click the mouse when the on-screen signal turns green, directly measuring visual-to-motor reaction latency.
Task
Metric
π0.5 H50E50
π0.5 H50E10
π0.5 H10E10
Async+RTC
REACT w/o DD
REACT
Human Avg.
Pour Rice (ARX X5)
Success (%) ↑
3
33
0
0
57
63
–
Reaction Game (Franka R3)
Latency (ms) ↓
1505 ± 502
768 ± 113
772 ± 119
747 ± 123
754 ± 127
734 ± 94
705 ± 73
Table 3: Specialized task results. Pour Rice reports success rate (%) (higher is better). Reaction Game reports mean ± standard deviation latency in milliseconds (lower is better).
Figure 5: Training efficiency test. (a) Training loss over optimization steps. (b) Average closed-loop success rate as a function of training steps; shaded bands show the task-wise range.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Over-pouring under long open-loop execution in Pour Rice. Consecutive frames from the H50E50 baseline show the policy continuing the pouring motion even after the rice level exceeds the target mark, illustrating how stale open-loop execution can overshoot the target state before the next observation update.
Figure 7: Trajectory response in Bowl Stacking by Size when the bowl is suddenly moved during grasping. Left: the original π0.5 full-chunk policy reacts slowly because it continues executing a stale action chunk. Middle: the shortened action-chunk baseline updates more frequently but produces discontinuous stop-and-replan motion. Right: REACT w/o DD adapts to the moved bowl while maintaining a smoother approach trajectory.
Figure 8: Repeated cap-contact retries under shortened execution in Bottle Cap Unscrewing. Representative frames from the H50E10 baseline show the policy repeatedly returning to the cap-contact region instead of sustaining the continuous grasp–twist–release sequence required to open the bottle.
Figure 9: Dual Decoupling timing diagram. Sensing, VLM encoding, DiT denoising, and action execution run as separate streams. The scheduler selects observations every ΔS=S/fcam from the camera stream, and parallel VLM workers encode them into a latest-ready embedding cache; each update consumes the freshest completed embedding and the rolling action buffer, then commits the next S executable actions. The architecture removes both the execution wait and the VLM wait from the DiT critical path: camera-rate input occurs in the S=1 setting, while output action throughput reaches the camera cadence whenever the VLM pool and DiT keep up.
Method
Φin (obs/sec)
Φout (act/sec)
Synchronous
Tinfer+He/fctrl1
Tinfer+He/fctrlHe
Sync + Accel ( α )
αTinfer+He/fctrl1
αTinfer+He/fctrlHe
Asynchronous
max(Tinfer,He/fctrl)1
max(Tinfer,He/fctrl)He
REACT w/o DD
TVLM+TDiT+S/fctrl1
TVLM+TDiT+S/fctrlS
REACT
min(Sfcam,TDiT1) , if RVLM≥fcam/S
min(fcam,TDiTS)
Appendix
Table 4: Throughput formulas for different inference schemes. Tinfer=TVLM+N⋅TDiT .
Method
He or S
Tcycle (ms)
Φin (Hz)
Φout (Hz)
Synchronous
He=50
1734
0.58
28.8
Synchronous
He=25
900
1.11
27.8
Synchronous
He=10
400
2.50
25.0
Synchronous
He=5
234
4.28
21.4
Sync + Accel ( α=0.5 )
He=50
1700
0.59
29.4
Sync + Accel ( α=0.5 )
He=25
867
1.15
28.8
Appendix
Table 5: Throughput comparison based on π0.5 ( H=50 , fcam=30 Hz, fctrl=30 Hz, TVLM=37 ms, TDiT=3 ms, N=10 ).
Figure 10: Hardware platforms. Representative views of the Acone/ARX X5 bimanual platform and the Franka Research 3 dual-arm platform used in the real-robot evaluation.
Hyperparameter
Value
Peak learning rate
2.5×10−5
Final decay learning rate
2.5×10−6
LR scheduler
Cosine decay with linear warmup
Warmup fraction
10% of the total training schedule
Optimizer
AdamW
AdamW β1
0.9
Appendix
Table 6: Training hyperparameters used for simulation and real-robot policies.
Policy
Latency (ms) ↓
Speedup
Notes
Standard π0.5
66.04
1.00 ×
VLM embedding + 10 denoising steps.
REACT w/o DD
39.96
1.65 ×
VLM embedding + single denoising step.
Appendix
Table 7: Real-robot per-update inference latency. Mean synchronous policy-update latency measured with torch.compile enabled.
Setting
Task
π0.5 H50E50
π0.5 H50E10
π0.5 H10E10
Async+RTC
REACT w/o DD
REACT
Simulation
beat_block_hammer
463.4
996.7
–
1755.0
664.0
775.9
click_bell
502.4
1087.2
1341.8
1946.4
690.8
712.0
move_playingcard_away
814.3
2012.7
1784.5
1832.2
1165.8
1341.8
pick_diverse_bottles
857.7
1391.2
–
1688.0
1253.4
1621.8
press_stapler
567.9
1295.5
1425.7
3168.2
1136.4
1204.7
rotate_qrcode
965.5
1663.9
1638.8
1944.5
1296.1
1422.3
Appendix
Table 8: Full trajectory smoothness results. Mean jerk norm of executed joint trajectories for each task and platform; lower indicates smoother motion. Best in bold , second best underlined .
Figure 11: Ablation on block length and block count. Mean RoboTwin clean success rate for each S×K configuration; hatched cells indicate configurations not evaluated.
Figure 12: Per-task ablation heatmaps. Success rate for each of the seven RoboTwin tasks as a function of block length S and block count K . Hatched cells indicate configurations not evaluated. The color scale is shared across all panels (0–100%).
Figure 13: OOD open-loop action prediction. Open-loop action-prediction MSE on the demo_randomized OOD dataset for three representative RoboTwin tasks.
Figure 14: Qualitative open-loop predictions on Move Playing Card Away. Two visualizations of the same demo_randomized example show how the predicted action sequence compares with the demonstration trajectory under the baseline and REACT .