Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning
Authors: Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, +7 more
Organizations: Beihang University · The Chinese University of Hong Kong · PKU-Psibot Lab · Tsinghua University · Zhongguancun Laboratory · Li Auto Inc. · Peking University
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both π0 and π0.5 backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by +23.1%, +16.4%, and +44% on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.
Figures & tables
Figure 1: The three-stage recipe for reinforcement fine-tuning multi-agent VLAs. We first perform Initialization-Aware Data Collection to broaden coverage of environment conditions, followed by Offline Credit-Filtered Tuning to decompose agent-wise advantages and reinforce high-contribution behaviors. Finally, we extend DSRL to the multi-agent setting and perform Online Latent-Space Reinforcement Learning to enable further improvement through trial and error.
Figure 2: Success rate (%, ↑ ) under 8 different simulation tasks from RoboTwin and RoboFactory with backbone models π0 and π0.5 .
Task
CHORUS
+ Offline RL w/o IA
+Offline RL +Online RL w/o IA
+Offline RL
+Offline RL +Online RL
Lift Pot
24
30
46
42
50
Handover Mic
43
46
63
52
73
Three Robots Stack Cubes
49
57
58
61
65
Take Photo
55
68
79
69
88
Average
42.8
50.3
61.5
56.0
69.0
Table 1: Ablation study on the initialization-aware data collection. We report success rate (%, ↑ ).
Figure 3: Visualization of success rates under different environment initializations. Initialization-Aware Data Collection identifies and compensates for challenging configurations in which agents struggle to acquire successful behaviors through autonomous exploration.
Figure 4: Visualization of agent-wise credit assignments and the behaviors of them.
Task
Offline RL
+Flow-SDE
+Flow-Noise
+DSRL
Lift Pot
42
43
43
50
Handover Mic
52
56
62
73
Three Robots Stack Cubes
61
62
63
65
Take Photo
69
82
78
88
Average
56.0
60.1
61.5
69.0
Table 2: Comparison of different online RL methods. We report success rate (%, ↑ ).
Figure 7
Figure 7: The success rate (%, ↑ ) and the overview under 3 different collaborative tasks on 2 Franka robots. Our training recipe effectively enhances the collaborative performance of multi-agent VLAs.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Batch size
32
Training steps
30,000
VLM
Full-parameter
Action expert
Full-parameter
Chunk horizon
50 for simulation and 20 for real-world
Appendix
Table 3: Training settings of the VLA policies and value functions.
Hyperparameter
Value
Hyperparameter
Value
Hyperparameter
Value
Batch size
16
Actor lr
1×10−4
Critic lr
3×10−4
Training episodes
200
Temperature lr
3×10−4
UTD ratio
1
Hidden dimensions
(128,128,128)
Latent dimension
50
Discount factor
0.999
Target update rate
0.005
Number of Q-functions
10
Critic reduction
Mean
CNN features
(32,32,32,32)
CNN strides
(2,1,1,1)
CNN padding
VALID
Encoder type
Small
Encoder normalization
Group
Bottleneck
True
Appendix
Table 4: Training hyperparameters of DSRL.
Hyperparameter
Value
Hyperparameter
Value
Hyperparameter
Value
Training episodes
200
Advantage type
embodied_grpo
Train expert only
True
Top-k
50
Loss type
embodied_grpo
Reward type
chunk_level
Parallel Num
8
PPO clip high
0.2
Separate models
True
Discount factor
0.99
PPO clip low
0.2
Action chunks
50
Chunk steps
16
PPO clip c
3.0
Flow steps
10
Eval chunk steps
16
Actor lr
5×10−6
Noise level
0.1 (Flow-SDE) 0.01 (Flow-Noise)
Appendix
Table 5: Training hyperparameters of Flow-SDE and Flow-Noise.
Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observations with linguistic instructions into desired robotic actions. Despite their advancements, VLAs often overlook explicit reasoning and learn the functional input-action mappings, omitting crucial logical steps, which are especially pronounced in interpretability and generalization for complex, long-horizon manipulation tasks. In this work, we propose ReFineVLA, a multimodal reasoning-aware framework that fine-tunes VLAs with teacher-guided reasons. We first augment robotic datasets with reasoning rationales generated by an expert teacher model, guiding VLA models to learn to reason about their actions. Then, we fine-tune pre-trained VLAs with the reasoning-enriched datasets with ReFineVLA, while maintaining the underlying generalization abilities and boosting reasoning capabilities. We also conduct attention map visualization to analyze the alignment among visual observation, linguistic prompts, and to-be-executed actions of ReFineVLA, reflecting the model is ability to focus on relevant tasks and actions. Through this additional step, we explore that ReFineVLA-trained models exhibit a meaningful agreement between vision-language and action domains, highlighting the enhanced multimodal understanding and generalization. Evaluated across a suite of simulated manipulation benchmarks on SimplerEnv with both WidowX and Google Robot tasks, ReFineVLA achieves state-of-the-art performance, in success rate over the second-best method on the both the WidowX benchmark and Google Robot Tasks.
Tuan Van Vo, Tan Q. Nguyen, Khang Nguyen +5
VinRobotics, Hanoi. Vietnam · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), UAE. · University of Texas at Arlington, Texas, USA +4
Vision-language-action models (VLAs) have emerged as a promising paradigm for general-purpose robot learning, with performance improving as models and datasets scale. Scaling robot data collection, however, remains challenging because data are naturally distributed across robots, tasks, and locations, making centralization costly or impractical. Federated learning offers a way to train on decentralized robot data, but applying it to VLAs requires accounting for heterogeneous robot client data distributions. We present Co-VLA, which applies consensus optimization using the Alternating Direction Method of Multipliers~(ADMM) to federated VLA training. We show that the same algorithm supports both full-model training and parameter-efficient fine-tuning with both fixed-rank and rank-adaptive adapters. The name Co-VLA reflects both consensus and collaboration: clients with different local robot datasets collaboratively train a shared model without sharing their data. Our experiments demonstrate that Co-VLA achieves performance comparable to centralized training in both full-model training and parameter-efficient fine-tuning settings. The project website and videos of our real-world experiments are available at https://embodiedvision.github.io/co-vla/.
Haolong Li, Guner Dilsad Er, Michael Muehlebach +1
Intelligent Perception in Technical Systems, University of Augsburg, Germany · Max Planck Institute for Intelligent Systems, Germany
Vision-language-action (VLA) models can learn to perform diverse manipulation skills "out of the box," but achieving the precision and speed that real-world tasks demand requires further fine-tuning -- for example, via reinforcement learning (RL). We introduce a lightweight method that enables sample-efficient online RL fine-tuning of pretrained VLAs using just a few hours of real-world practice. We (1) adapt the VLA to expose an "RL token," a compact readout representation that preserves task-relevant pretrained knowledge while serving as an efficient interface for online RL, and (2) train a small actor-critic head on this RL token to refine the actions, while anchoring the learned policy to the VLA. Online RL with the RL token (RLT) makes it possible to fine-tune even large VLAs with RL quickly and efficiently. Across four real-robot tasks (screw installation, zip tie fastening, charger insertion, and Ethernet insertion), RLT improves the speed on the hardest part of the task by up to 3x and raises success rates significantly within minutes to a few hours of practice. It can even surpass the speed of human teleoperation on some of the tasks.
Charles Xu, Jost Tobias Springenberg, Michael Equi +4