RoboFL: Federated Expert Assembly for World Action Models
Authors: Rongyu Zhang, Ruizhi Fan, Yunfan Lou, Hengyu Fang, Shenli Zheng, Chenrui Wu, Yili Jin, Li Du, +3 more
Organizations: Nanjing University · Hong Kong University of Science and Technology · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient fine-tuning, avoiding the exchange of full-model updates. However, federating these adapters is nontrivial, as naive aggregation can entangle incompatible updates, while incorporating MoE-style routing into federated aggregation may dilute specialization and destabilize expert selection. We present RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning. MoSAIC directly installs locally trained LoRA adapters as the expert branches of a server MoE. Server-side routers learn token assignments over these prior-informed branches while jointly refining routing and expert parameters. Foresight-to-Action Routing Distillation (FARD) aligns routing across the model's three paths, while Path-Consensus Expert Aggregation (PCEA) converts complete expert updates into a compact global adapter for personalized redistribution. Experiments on RoboTwin 2.0, RLBench, and a real-world Franka robot arm show the superiority of RoboFL with structured expert assembly, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.
Figures & tables
Figure 1: From federated parameter averaging to prior-informed expert assembly. Left: Aggregating independently trained LoRA-MoE models can dilute local specialization and introduce inconsistent routing across heterogeneous clients. Middle: RoboFL organizes institutions into task silos and directly installs their task-trained LoRA into the expert branches of a server MoE. Right: RoboFL is evaluated in the RoboTwin 2.0 and RLBench simulator, and on a real-world Franka robot.
Figure 2: Overview of RoboFL . Left: Modality-specific tokens interact through unified masked self-attention, while FARD distills detached understanding-generation consensus into action routing. Right: MoSAIC installs task-trained LoRA adapters as server MoE experts and jointly refines routers and experts. PCEA aggregates complete expert updates using three-path consensus to form a rank-constrained global adapter, which is blended with each server-refined client expert.
Method
beat ham.
rank size
hand blo.
hang mug
lift pot
move can
move sta.
pick div.
pick dual
place L
place R
bread ski.
can basket
CENTRALIZED TRAINING
InternVLA
73
78
62
27
32
76
53
68
74
85
84
79
60
Motus
82
86
60
20
40
86
48
70
80
85
80
80
69
FEDERATED LEARNING (LoRA/MoE)
FedAvg
75
74
52
24
33
67
39
69
77
79
79
74
72
FedMoE
69
71
42
21
34
57
34
61
70
84
80
66
64
Table 1: RoboTwin 2.0 success rates on a 25-task challenging subset under random condition (selected as tasks where FedAvg < 80%). Overall reports performance on all 50 tasks.
Method
InternVLA
Motus
ForgeVLA
FedVLA
RoboFL
Type
Central-LoRA
FL-LoRA
FL-MoE
LoRA → MoE
Success
52.75
52.50
34.25
31.00
41.25
Table 2: Average success rates (%) for eight RLBench tasks with 4-expert MoSAIC for RoboFL .
Method
Adapters c / s
Params (M)
Comm. MiB
Mem. GiB
Success (%)
FedAvg
1 / 1
40.85
311.63
8.33
78.92
FedMoE
4 / 8
158.11
1,206.28
11.44
75.12
ForgeVLA
1 / 1
40.85
312.06
8.33
80.70
FedVLA
8 / 8
312.93
2,362.85
15.59
79.32
RoboFL
1 / 8
40.85
311.63
8.33
83.12
Table 3: Client resources and logical communication on RoboTwin 2.0. Communication combines upload and download per participating client per round.
Figure 3: Success-threshold coverage on RoboTwin. Bars count tasks satisfying each success-rate criterion across 50 tasks with 100 trials per task.
Figure 4: Route-to-action coupling under instruction swaps and direct action-router interventions. The inset magnifies the direct intervention regime.
Method
Type
adjust bottle
stamp seal
stack cups
screw bottle
pour water
test tube
Mean ± Avg. SD
InternVLA
Centralized-LoRA
63.33
53.33
41.67
31.67
48.33
43.33
46.94±7.80
ForgeVLA
FL-LoRA
41.67
31.67
46.67
28.33
31.67
33.33
35.56±7.03
FedVLA
FL-MoE
43.33
41.67
33.33
13.33
40.00
30.00
33.61±5.60
RoboFL
FL-LoRA → MoE
68.33
66.67
68.33
36.67
56.67
58.33
59.17±7.61
Table 4: Real-world applications with six single-arm Franka tasks SR (%) with three replicates.
Figure 5: Qualitative successful real-world Franka rollouts for each of six single-arm Franka tasks.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
RoboTwin 2.0
RLBench
Franka
Optimization and adapters
Pretrained checkpoint
InternVLA-A1-3B
Clients / experts
8/8
4/4
4/4
Routing top- k
4
2
2
Communi. rounds
100
100
100
Local steps / round
500
200
300
Appendix
Table 5: Training and evaluation settings for the simulation and real-world scenarios. Execution horizon is specified independently of prediction length.
Table 6: Task assignments in the explicit eight-client task-silo implementation. Counts refer to tasks, not episodes or training frames.
Method
adjust bot.
rank RGB
click alarm
click bell
dump bin
grab roller
hand mic
move pill
move card
open laptop
open micro.
bread bas.
burger fries
CENTRALIZED TRAINING
InternVLA
99
91
85
91
98
100
95
87
100
98
91
87
99
Motus
100
94
83
90
91
100
81
91
97
92
80
81
93
FEDERATED LEARNING (LoRA/MoE)
FedAvg
99
91
84
93
95
100
89
83
98
90
84
82
98
FedMoE
100
91
76
91
94
100
82
80
95
88
65
80
97
Appendix
Table 7: RoboTwin 2.0 task success rates (%) on the complementary 25 tasks not shown in Table 1 . Best and second-best reported results are boldfaced and italicized.
Figure 6: Franka single-arm experimental setup. The platform comprises a Franka Research 3, a Robotiq 2F-85 gripper, and two RealSense D435i cameras mounted on the wrist and an external stand, respectively. The tabletop contains the objects and fixtures used in the manipulation tasks.
Configuration
Overall
FedAvg (averaged single adapter)
78.92
FedMoE (client-side routed MoE)
75.12
Vanilla MoSAIC (task-trained experts + routing)
81.94
+ FARD
82.42
+ FARD + PCEA ( RoboFL )
83.12
Appendix
Table 10: Ablation of MoSAIC components on RoboTwin 2.0. All rows report overall success rates (%) on all 50 tasks with 100 trials per task. The three MoSAIC rows use the same federated pipeline and differ only in the server mechanism; FedAvg and FedMoE are external references.
Method
put rubbish
seat down
draw umbrella
close laptop
sweep dustpan
close fridge
close box
phone on base
Overall
CENTRALIZED TRAINING
InternVLA
60
18
34
62
66
46
86
50
52.75
Motus
44
22
36
58
72
58
92
38
52.50
FEDERATED LEARNING (LoRA/MoE)
FedAvg
0
18
22
4
14
72
92
0
27.75
FedMoE
0
12
16
14
22
64
70
8
25.75
Appendix
Table 11: RLBench task success rates (%) on the 8-task single-view suite. Overall is the equally weighted mean across the eight tasks. Best and second-best results are boldfaced and italicized.
Figure 7: Qualitative RoboTwin 2.0 rollouts (part 1 of 2). Each row shows two tasks. Every task is labeled with its name and contributes four chronologically ordered keyframes from one successful rollout, illustrating approach, contact or grasp, and execution of the intended state change. Frames are exported at their native resolution without cropping or enhancement.
Figure 8: Qualitative RoboTwin 2.0 rollouts (part 2 of 2). Each row shows two tasks. Every task is labeled with its name and contributes four chronologically ordered keyframes from one successful rollout, illustrating approach, contact or grasp, and execution of the intended state change. Frames are exported at their native resolution without cropping or enhancement.
Figure 9: Qualitative RLBench rollouts. Each row shows two tasks, each labeled with its name and the number of successful episodes out of 50, and includes up to four chronologically ordered keyframes from one successful rollout. All rollouts use the single-view interface with the front camera at variation 0 and the eight-action chunk-queue protocol. The displayed episode for each task was selected for visual clarity rather than being the longest or most representative rollout.
Figure 10: Failure cases across benchmarks. RoboTwin 2.0 failures (top) and RLBench failures (bottom) contribute four chronologically ordered keyframes from one failed rollout: initial scene, first meaningful attempt, visible deviation, and final failing state. Because observations are recorded before actions and the video ends at the failure, the last frame shows the episode-terminated state.
Figure 11: Franka failure cases. Failed rollouts for six Franka tasks, with each row contributing four chronologically ordered keyframes (initial scene, first attempt, visible deviation, final state).