Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.
Figures & tables
Fig. 1: In current practice (left), each site downloads a released VLA checkpoint, and the demonstrations it collects afterwards go unused. decent-vla (right) lets each site fine-tune the shared policy on its private data and exchange only model updates, so the policy keeps improving while data stays local.
Fig. 5: Three physical sites (A–C): the overhead and wrist camera views that form the policy’s input. Heterogeneity is both visual and physical, the latter from differing hardware calibration (Fig. 9 ).
Fig. 6: The simulated environment of SO-101. Left: the overhead and wrist views the policy receives. Right: two rollouts of one π0.5 checkpoint, one completing the stack and one not, with each stage shown as a composite of the five frames before it triggers.
Training
Spatial
Object
Goal
Long
Overall
π0.5 (3.2 B) [ 1 ]
Centralized
97.0
99.0
97.0
95.0
97.00
Federated
92.0
100.0
99.0
99.0
97.50
GR00T N1.7 (3.1 B) [ 3 ]
Centralized
97.0
97.0
99.0
94.0
96.75
Federated
91.0
99.0
95.0
79.0
91.00
TABLE I: Federated versus centralized fine-tuning: LIBERO success rate (%); best per column in bold.
Update parameterization
Params
Payload
Success (%)
π0.5
Full fine-tuning
3.62 B
7.24 GB
97.50
LoRA
0.12 B
0.24 GB
97.50
GR00T N1.7
Action head (full)
1.62 B
3.24 GB
91.00
Action head (LoRA)
0.39 B
0.77 GB
87.00
TABLE II: Full fine-tuning versus LoRA: payload and LIBERO success rate (%); best per policy in bold.
Federated scope
Shared
Personal
Payload
Personalized success (%)
π0.5
Action head
0.69 B
2.92 B
1.39 GB
95.25
Backbone
2.92 B
0.69 B
5.85 GB
97.25
GR00T N1.7
Action head
1.62 B
1.52 B
3.24 GB
86.25
Backbone
1.52 B
1.62 B
3.05 GB
71.25
TABLE III: Partial federation: personalized LIBERO success rate (%); best per policy in bold. Only the listed scope is federated, the remainder stays client-personal.
Algorithm
Spatial
Object
Goal
Long
Overall
FedAvg [ 4 ]
92.0
100.0
99.0
99.0
97.50
FedAdam [ 17 ]
91.0
100.0
99.0
95.0
96.25
FedProx [ 15 ]
89.0
99.0
95.0
94.0
94.25
TABLE IV: Aggregation algorithms on π0.5 full fine-tuning: LIBERO success rate (%); best per column in bold.
Training
Base
Language
Object
Position
Task
Overall
π0.5
Centralized
97.00
97.50
87.50
22.50
4.50
53.00
Federated
97.50
97.00
83.50
43.25
4.75
57.13
GR00T N1.7
Centralized
96.75
93.25
79.00
2.00
5.00
44.81
Federated
91.00
92.25
73.75
8.75
4.00
44.69
TABLE V: Robustness to distribution shift: LIBERO-PRO success rate (%); best per column in bold.