Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.
Figures & tables
Fig. 1: In current practice (left), each site downloads a released VLA checkpoint, and the demonstrations it collects afterwards go unused. decent-vla (right) lets each site fine-tune the shared policy on its private data and exchange only model updates, so the policy keeps improving while data stays local.
Fig. 5: Three physical sites (A–C): the overhead and wrist camera views that form the policy’s input. Heterogeneity is both visual and physical, the latter from differing hardware calibration (Fig. 9 ).
Fig. 6: The simulated environment of SO-101. Left: the overhead and wrist views the policy receives. Right: two rollouts of one π0.5 checkpoint, one completing the stack and one not, with each stage shown as a composite of the five frames before it triggers.
Training
Spatial
Object
Goal
Long
Overall
π0.5 (3.2 B) [ 1 ]
Centralized
97.0
99.0
97.0
95.0
97.00
Federated
92.0
100.0
99.0
99.0
97.50
GR00T N1.7 (3.1 B) [ 3 ]
Centralized
97.0
97.0
99.0
94.0
96.75
Federated
91.0
99.0
95.0
79.0
91.00
TABLE I: Federated versus centralized fine-tuning: LIBERO success rate (%); best per column in bold.
Update parameterization
Params
Payload
Success (%)
π0.5
Full fine-tuning
3.62 B
7.24 GB
97.50
LoRA
0.12 B
0.24 GB
97.50
GR00T N1.7
Action head (full)
1.62 B
3.24 GB
91.00
Action head (LoRA)
0.39 B
0.77 GB
87.00
TABLE II: Full fine-tuning versus LoRA: payload and LIBERO success rate (%); best per policy in bold.
Federated scope
Shared
Personal
Payload
Personalized success (%)
π0.5
Action head
0.69 B
2.92 B
1.39 GB
95.25
Backbone
2.92 B
0.69 B
5.85 GB
97.25
GR00T N1.7
Action head
1.62 B
1.52 B
3.24 GB
86.25
Backbone
1.52 B
1.62 B
3.05 GB
71.25
TABLE III: Partial federation: personalized LIBERO success rate (%); best per policy in bold. Only the listed scope is federated, the remainder stays client-personal.
Algorithm
Spatial
Object
Goal
Long
Overall
FedAvg [ 4 ]
92.0
100.0
99.0
99.0
97.50
FedAdam [ 17 ]
91.0
100.0
99.0
95.0
96.25
FedProx [ 15 ]
89.0
99.0
95.0
94.0
94.25
TABLE IV: Aggregation algorithms on π0.5 full fine-tuning: LIBERO success rate (%); best per column in bold.
Training
Base
Language
Object
Position
Task
Overall
π0.5
Centralized
97.00
97.50
87.50
22.50
4.50
53.00
Federated
97.50
97.00
83.50
43.25
4.75
57.13
GR00T N1.7
Centralized
96.75
93.25
79.00
2.00
5.00
44.81
Federated
91.00
92.25
73.75
8.75
4.00
44.69
TABLE V: Robustness to distribution shift: LIBERO-PRO success rate (%); best per column in bold.
Vision-language-action models (VLAs) have emerged as a promising paradigm for general-purpose robot learning, with performance improving as models and datasets scale. Scaling robot data collection, however, remains challenging because data are naturally distributed across robots, tasks, and locations, making centralization costly or impractical. Federated learning offers a way to train on decentralized robot data, but applying it to VLAs requires accounting for heterogeneous robot client data distributions. We present Co-VLA, which applies consensus optimization using the Alternating Direction Method of Multipliers~(ADMM) to federated VLA training. We show that the same algorithm supports both full-model training and parameter-efficient fine-tuning with both fixed-rank and rank-adaptive adapters. The name Co-VLA reflects both consensus and collaboration: clients with different local robot datasets collaboratively train a shared model without sharing their data. Our experiments demonstrate that Co-VLA achieves performance comparable to centralized training in both full-model training and parameter-efficient fine-tuning settings. The project website and videos of our real-world experiments are available at https://embodiedvision.github.io/co-vla/.
Haolong Li, Guner Dilsad Er, Michael Muehlebach +1
Intelligent Perception in Technical Systems, University of Augsburg, Germany · Max Planck Institute for Intelligent Systems, Germany
Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring annotated training data. Fortunately, vision-equipped robots deployed across various domains already produce abundant vision-action pairs that can be leveraged to scale up VLA training more efficiently. However, these raw data cannot be centrally aggregated due to various constraints and also exhibit severe heterogeneity. To address these challenges, in this paper, we propose ForgeVLA, a federated VLA training framework that learns VLA models from distributed vision-action pairs without centralizing raw data or requiring manual annotations. Specifically, each client in ForgeVLA is equipped with an embodied instruction classifier that maps vision-action pairs to a predefined instruction set, recovering the missing language modality and forming complete vision-language-action triplets. Beyond triplet construction, we also identify vision-language feature collapse as a critical challenge that has been largely overlooked in prior federated VLA research. To mitigate this issue, ForgeVLA combines a client-side contrastive planning loss with a server-side adaptive aggregation strategy to learn task-discriminative representations efficiently. Extensive experiments across multiple benchmarks show that ForgeVLA significantly outperforms other baselines, and ablation studies further validate the contribution of each component.
Yuhao Zhou, Yunpeng Zhu, Yang Zhou +7
Sichuan University · Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education · Zhejiang University +2
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both π0 and π0.5 backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by +23.1%, +16.4%, and +44% on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.
Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu +12
Beihang University · The Chinese University of Hong Kong · PKU-Psibot Lab +4