cs.ROSep 28, 2026

RoboFL: Federated Expert Assembly for World Action Models

Authors: Rongyu Zhang, Ruizhi Fan, Yunfan Lou, Hengyu Fang, Shenli Zheng, Chenrui Wu, Yili Jin, Li Du, +3 more

Organizations: Nanjing University · Hong Kong University of Science and Technology · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University

Abstract

Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient fine-tuning, avoiding the exchange of full-model updates. However, federating these adapters is nontrivial, as naive aggregation can entangle incompatible updates, while incorporating MoE-style routing into federated aggregation may dilute specialization and destabilize expert selection. We present RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning. MoSAIC directly installs locally trained LoRA adapters as the expert branches of a server MoE. Server-side routers learn token assignments over these prior-informed branches while jointly refining routing and expert parameters. Foresight-to-Action Routing Distillation (FARD) aligns routing across the model's three paths, while Path-Consensus Expert Aggregation (PCEA) converts complete expert updates into a compact global adapter for personalized redistribution. Experiments on RoboTwin 2.0, RLBench, and a real-world Franka robot arm show the superiority of RoboFL with structured expert assembly, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Co-VLA: Consensus-based Federated Training for Vision-Language-Action Models

    Sep 17, 2026Haolong Li, Guner Dilsad Er, Michael Muehlebach +1Diffusion-Based Vision-Language-ActionsScalable Robot Learning

  2. ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations

    May 8, 2026Yuhao Zhou, Yunpeng Zhu, Yang Zhou +7Diffusion-Based Vision-Language-ActionsContrastive Loss

  3. An Empirical Study and Open Testbed for Federated Fine-Tuning of Vision-Language-Action Models

    Sep 19, 2026Zhekai Duan, Kevin Ziyang Xie, Xinyu Tan +5Libero Manipulation BenchmarkModel Fine-Tuning