RoboFL: Federated Expert Assembly for World Action Models
Organizations: Nanjing University · Hong Kong University of Science and Technology · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Abstract
Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient fine-tuning, avoiding the exchange of full-model updates. However, federating these adapters is nontrivial, as naive aggregation can entangle incompatible updates, while incorporating MoE-style routing into federated aggregation may dilute specialization and destabilize expert selection. We present RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning. MoSAIC directly installs locally trained LoRA adapters as the expert branches of a server MoE. Server-side routers learn token assignments over these prior-informed branches while jointly refining routing and expert parameters. Foresight-to-Action Routing Distillation (FARD) aligns routing across the model's three paths, while Path-Consensus Expert Aggregation (PCEA) converts complete expert updates into a compact global adapter for personalized redistribution. Experiments on RoboTwin 2.0, RLBench, and a real-world Franka robot arm show the superiority of RoboFL with structured expert assembly, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.
Figures & tables
| Method | beat ham. | rank size | hand blo. | hang mug | lift pot | move can | move sta. | pick div. | pick dual | place L | place R | bread ski. | can basket |
| CENTRALIZED TRAINING | |||||||||||||
| InternVLA | 73 | 78 | 62 | 27 | 32 | 76 | 53 | 68 | 74 | 85 | 84 | 79 | 60 |
| Motus | 82 | 86 | 60 | 20 | 40 | 86 | 48 | 70 | 80 | 85 | 80 | 80 | 69 |
| FEDERATED LEARNING (LoRA/MoE) | |||||||||||||
| FedAvg | 75 | 74 | 52 | 24 | 33 | 67 | 39 | 69 | 77 | 79 | 79 | 74 | 72 |
| FedMoE | 69 | 71 | 42 | 21 | 34 | 57 | 34 | 61 | 70 | 84 | 80 | 66 | 64 |
| Method | InternVLA | Motus | ForgeVLA | FedVLA | RoboFL |
| Type | Central-LoRA | FL-LoRA | FL-MoE | LoRA MoE | |
| Success | 52.75 | 52.50 | 34.25 | 31.00 | 41.25 |
| Method | Adapters c / s | Params (M) | Comm. MiB | Mem. GiB | Success (%) |
| FedAvg | 1 / 1 | 40.85 | 311.63 | 8.33 | 78.92 |
| FedMoE | 4 / 8 | 158.11 | 1,206.28 | 11.44 | 75.12 |
| ForgeVLA | 1 / 1 | 40.85 | 312.06 | 8.33 | 80.70 |
| FedVLA | 8 / 8 | 312.93 | 2,362.85 | 15.59 | 79.32 |
| RoboFL | 1 / 8 | 40.85 | 311.63 | 8.33 | 83.12 |
| Method | Type | adjust bottle | stamp seal | stack cups | screw bottle | pour water | test tube | Mean Avg. SD |
| InternVLA | Centralized-LoRA | 63.33 | 53.33 | 41.67 | 31.67 | 48.33 | 43.33 | |
| ForgeVLA | FL-LoRA | 41.67 | 31.67 | 46.67 | 28.33 | 31.67 | 33.33 | |
| FedVLA | FL-MoE | 43.33 | 41.67 | 33.33 | 13.33 | 40.00 | 30.00 | |
| RoboFL | FL-LoRA MoE | 68.33 | 66.67 | 68.33 | 36.67 | 56.67 | 58.33 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | RoboTwin 2.0 | RLBench | Franka |
| Optimization and adapters | |||
| Pretrained checkpoint | InternVLA-A1-3B | ||
| Clients / experts | |||
| Routing top- | |||
| Communi. rounds | |||
| Local steps / round | |||
| Partition | Count | Tasks |
| Client 0 | 5 | grab_roller , pick_diverse_bottles , pick_dual_bottles , put_bottles_dustbin , put_object_cabinet |
| Client 1 | 7 | place_object_basket , place_bread_basket , place_bread_skillet , place_can_basket , place_cans_plasticbox , place_container_plate , place_empty_cup |
| Client 2 | 8 | place_a2b_left , place_a2b_right , place_fan , place_mouse_pad , place_object_scale , place_object_stand , place_phone_stand , place_shoe |
| Client 3 | 6 | stack_blocks_three , stack_blocks_two , stack_bowls_three , stack_bowls_two , blocks_ranking_rgb , blocks_ranking_size |
| Client 4 | 5 | adjust_bottle , lift_pot , move_pillbottle_pad , move_playingcard_away , move_stapler_pad |
| Client 5 | 3 | beat_block_hammer , press_stapler , stamp_seal |
| Method | adjust bot. | rank RGB | click alarm | click bell | dump bin | grab roller | hand mic | move pill | move card | open laptop | open micro. | bread bas. | burger fries |
| CENTRALIZED TRAINING | |||||||||||||
| InternVLA | 99 | 91 | 85 | 91 | 98 | 100 | 95 | 87 | 100 | 98 | 91 | 87 | 99 |
| Motus | 100 | 94 | 83 | 90 | 91 | 100 | 81 | 91 | 97 | 92 | 80 | 81 | 93 |
| FEDERATED LEARNING (LoRA/MoE) | |||||||||||||
| FedAvg | 99 | 91 | 84 | 93 | 95 | 100 | 89 | 83 | 98 | 90 | 84 | 82 | 98 |
| FedMoE | 100 | 91 | 76 | 91 | 94 | 100 | 82 | 80 | 95 | 88 | 65 | 80 | 97 |
| Partition | Count | Tasks |
| Client 0 | 1 | take_umbrella_out_of_umbrella_stand |
| Client 1 | 2 | close_fridge , toilet_seat_down |
| Client 2 | 2 | close_laptop_lid , close_box |
| Client 3 | 1 | sweep_to_dustpan |
| Server | 2 | put_rubbish_in_bin , phone_on_base |
| Partition | Count | Tasks |
| Client 0 | 1 | stack_cups |
| Client 1 | 1 | screw_bottle |
| Client 2 | 1 | stamp_seal |
| Client 3 | 1 | test_tube |
| Server | 2 | adjust_bottle , pour_water |
| Configuration | Overall |
| FedAvg (averaged single adapter) | 78.92 |
| FedMoE (client-side routed MoE) | 75.12 |
| Vanilla MoSAIC (task-trained experts + routing) | 81.94 |
| FARD | 82.42 |
| FARD PCEA ( RoboFL ) | 83.12 |
| Method | put rubbish | seat down | draw umbrella | close laptop | sweep dustpan | close fridge | close box | phone on base | Overall |
| CENTRALIZED TRAINING | |||||||||
| InternVLA | 60 | 18 | 34 | 62 | 66 | 46 | 86 | 50 | 52.75 |
| Motus | 44 | 22 | 36 | 58 | 72 | 58 | 92 | 38 | 52.50 |
| FEDERATED LEARNING (LoRA/MoE) | |||||||||
| FedAvg | 0 | 18 | 22 | 4 | 14 | 72 | 92 | 0 | 27.75 |
| FedMoE | 0 | 12 | 16 | 14 | 22 | 64 | 70 | 8 | 25.75 |