cs.ROSep 29, 2026

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

Authors: Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, +7 more

Organizations: Beihang University · The Chinese University of Hong Kong · PKU-Psibot Lab · Tsinghua University · Zhongguancun Laboratory · Li Auto Inc. · Peking University

Abstract

We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both π0π_0 and π0.5π_{0.5} backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by +23.1%+23.1\%, +16.4%+16.4\%, and +44%+44\% on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning

    Apr 20, 2026Tuan Van Vo, Tan Q. Nguyen, Khang Nguyen +5Diffusion-Based Vision-Language-ActionsMultimodal Reasoning

  2. Co-VLA: Consensus-based Federated Training for Vision-Language-Action Models

    Sep 17, 2026Haolong Li, Guner Dilsad Er, Michael Muehlebach +1Diffusion-Based Vision-Language-ActionsScalable Robot Learning

  3. RL Token: Bootstrapping Online RL with Vision-Language-Action Models

    Apr 24, 2026Charles Xu, Jost Tobias Springenberg, Michael Equi +4Diffusion-Based Vision-Language-ActionsReinforcement Fine-Tuning