Federated Reinforcement Learning (FRL) enables collaborative policy learning across distributed agents with heterogeneous environments. While recent methods based on variance reduction, divergence penalization, and momentum optimization improve FRL under heterogeneous settings, they still primarily synchronize policy or value-network parameters and do not explicitly address distributional mismatch among heterogeneous clients. Therefore, we propose \textbf{FedGuide}, a FRL framework that uses diffusion priors as behavior models to provide personalized data supported distributions for heterogeneous local policy learning. Instead of directly averaging local policies, FedGuide aggregates those diffusion priors through Optimal-Transport Mixture-of-Experts (OT-MoE), preserving heterogeneous behavior modes in distribution space. It further develops a Distribution Correction Estimation (DICE) value baseline to provide low-variance, return-aware guidance for local policy improvement. Experiments across heterogeneous environments show that FedGuide outperforms representative FRL methods in client-average returns, final-round performance, and worst-round robustness, while maintaining stable learning under stronger heterogeneity.
Figures & tables
Figure 1: Method Visualization with 2 Clients Heterogeneous Settings. (a) Policy aggregation averages local optima x1∗,x2∗ into global policy x(t,H) , away from the global optimum x∗ . (b) OT-MoE aggregates diffusion priors in distribution space, forming shared prior support πˉD and personalized priors πˉD,1,πˉD,2 (pink). The prior regularizer R (yellow) pulls local policies toward personalized prior supports. (c) FedGuide further uses the DICE value baseline (green) for policy improvement, shown as straighter policies updates and smaller endpoint dispersion ( Variance↓ ).
Figure 2
Variants
Diffusion Prior
DICE Value
Client Policy
FG-A
OT-MoE
Local
Avg.
FG-P
OT-MoE
N/A
Local
FG
OT-MoE
Local
Local
Table 1: Comparison of FedGuide (FG) variants: FedGuide keeps the DICE value baseline Vi local, while FedGuide-P does not use Vϕ,i ( β=1 ). FedGuide-A uses Vϕ,i but periodically averages client policies at the server. All variants use OT-MoE aggregation for diffusion prior.
CVσT↓
Env
FG-A
FG-P
FG
Reacher
0.380
0.068
0.053
Hopper
0.402
0.164
0.155
Walker2D
0.208
0.179
0.127
HalfCheetah
0.628
0.235
0.241
MetaWorld10
0.364
0.272
0.260
Table 2: Normalized temporal volatility CVσT for FedGuide (FG) and variants, computed over the last 20 rounds as the across-seed mean of return standard deviation normalized by the absolute mean return. Smaller values indicate smoother late-stage returns and are consistent with, but do not directly measure, the variance-reduction effect in Lemma 2.
Figure 2: Left: Demonstrations of different heterogeneous FRL environments: Reacher, Walker2D, Hopper, HalfCheetah, and MetaWorld10 with 10 heterogeneous manipulation tasks. Right: Visualization of toy-case Bandit2D with 4 isotropic Gaussian reward peaks ( σ=0.2 ) evenly spaced on the unit circle, each client sampled from an overlapping 120∘ sector of the annulus r∈[0.7,1.3] centered on one peak. All clients in FedAvg collapse to one peak, while FedGuide (including its two variants) captures each peak; its global diffusion prior recovers the ring-shaped union of client supports through OT-MoE aggregation.
Figure 3: Client average return for 100 rounds at different heterogeneous FRL environments. We set client numbers to 10 in MetaWorld10 and keep 8 clients in other environments, evaluating on 5 random seeds with Behavior Cloning warm-up. Color regions indicate one standard deviation.
Figure 4: Client average return for FRL environments with stronger heterogeneity (“Hard”). We fixed client numbers at 8 and the same experimental settings with Fig. 3 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Client-level heterogeneity heatmaps for the main benchmark settings. Hopper, Walker2D, and HalfCheetah show normalized dynamics and reward metadata; Reacher shows normalized goal, action-noise, reward, and angle-shift metadata; MetaWorld10 shows task-semantic attributes.
Figure 6: Main-to-Hard metadata range expansion. For each metadata dimension, client values are normalized by the pooled Main and Hard range and plotted as side-by-side Main and Hard distributions.
Method
Heterogeneity Type
Aggregation Target
Alignment Space
Multi-Modal Support
Local Policy
Return-Aware Baseline
Baselines
FedAvg [ 23 ]
Data
Policy Params.
Parameter Space
✗
✗
✗
FedKL [ 44 ]
Data
Policy Params.
Policy Output
✗
✗
P
FedRL [ 13 ]
MDP / Dynamics
Actor + Critic
Parameter Space
✗
✗
P
FedSVRPG-M [ 37 ]
MDP / Dynamics
Policy Grad.
Gradient Space
✗
✗
P
Variants
FedGuide-P
Data / MDP / Task
Diffusion Prior
Distribution Space
✓
✓
✗
FedGuide-A
Data / MDP / Task
Prior + Policy
Distribution Space
✓
✗
✓
Appendix
Table 4: Comparison of the baselines and FedGuide variants used in our experiments. FedAvg aggregates policy parameters, FedKL regularizes local policies through KL penalties in the policy-output space, FedRL aggregates actor-critic models, and FedSVRPG-M aggregates variance-reduced policy-gradient updates. FedGuide instead aggregates diffusion behavior priors in distribution space while keeping each client’s policy local. FedGuide-P removes the DICE value baseline, and FedGuide-A additionally averages client policies. Here, ✓ denotes explicit support, ✗ denotes no explicit support, and P denotes partial or indirect support.
Method
Convergence
Communication Speedup
Local Updates Allowed
Distribution-Level Analysis
Baselines
FedAvg [ 23 ]
FL convergence
Empirical
✓
✗
FedKL [ 44 ]
Asymptotic
Not claimed
P
Policy-output KL
FedRL [ 13 ]
Finite to subopt.
Not claimed
✓
✗
FedSVRPG-M [ 37 ]
Variance-reduced PG
Claimed
✓
Policy-gradient correction
Variants
FedGuide-P
Prior-only ablation
Not isolated
✓
OT-MoE diffusion prior
FedGuide-A
Controlled avg. error
Not isolated
✓
OT-MoE prior + policy avg.
Appendix
Table 5: Theoretical and mechanism-level comparison of the baselines and FedGuide variants. FedAvg provides the standard federated policy averaging baseline, FedKL provides convergence analysis for KL-regularized policy-output alignment, FedRL analyzes federated RL under environment heterogeneity with convergence to a heterogeneity dependent suboptimal solution, and FedSVRPG-M uses momentum-based variance-reduced policy-gradient correction for efficient heterogeneous FRL. FedGuide keeps policies local and aggregates only diffusion-prior heads through OT-MoE; its analysis separates the effective stochastic variance, OT errors, prior or value surrogate errors, and local KL drift, while the DICE value baseline reduces the variance of local policy updates.
Parameter
Symbol
Bandit2D
Reacher
MuJoCo locomotion
MetaWorld10
Offline buffer size
∣Di∣
103
2×104
2×105 cap
5×103
Offline source
–
local modes
client rollouts
D4RL
scripted policy
Diffusion steps
K
Gaussian
1000
1000
1000
Prior width / horizon
dD,LD
Gaussian
64,64
64,64
64,64
Prior epochs / batch
ED,BD
closed form
40,512
40,512
40,512
Prior learning rate
ηD
–
10−4
10−4
10−4
Appendix
Table 6: Offline pretraining hyperparameters.
Parameter
Symbol
Bandit2D
Reacher
MuJoCo locomotion
MetaWorld10
Clients
N
4
8
8
10
Prior experts
M
4
8
8
8
Federated rounds
H
60
100
100
100
Rollout size / round
Troll
200
2048
4096
4096
Local PPO epochs
U
4
4
4
4
PPO mini-batch
B
64
64
64
64
Appendix
Table 7: Shared online hyperparameters for FedGuide variants.
Method
Prior aggregation
Value baseline
Policy aggregation
Key coefficients
FedGuide
OT-MoE on {ψi}i=1N
β=0.5
none
λ1=0.05 , λ2=0.5 , η=0.05
FedGuide-P
OT-MoE on {ψi}i=1N
β=1
none
λ1=0.05 , λ2=0.5 , η=0.05
FedGuide-A
OT-MoE on {ψi}i=1N
β=0.5
every 5 rounds
λ1=0.05 , λ2=0.5 , η=0.05
Appendix
Table 8: FedGuide variant hyperparameters.
Parameter
Symbol
FedAvg
FedKL
FedRL
FedSVRPG-M
Local optimizer
–
PPO
PPO
DDPG
PPO/SVRPG
Aggregated parameters
–
θi
θi
θiμ,θiQ
Δθi
Global KL weight
λg
0
0.1
–
–
Local KL weight
λl
0
0.05
–
–
Local PPO epochs
U
10
10
–
4
Mini-batch size
B
64
64
64
64
Appendix
Table 9: Baseline hyperparameters.
Figure 7: Bandit2D sensitivity to the expert count M , the Sinkhorn temperature η , and the prior weight λ2 .
Figure 8: Weak-overlap stress test. Client 5 has a true target mode at the origin, while the four experts contain only ring-shaped behavior support. OT-MoE routes Client 5 to a mixture of the existing experts, keeping its personalized prior on the ring rather than creating the unseen center mode.
Figure 9: FedGuide-family qualitative rollouts on Reacher, visualized with motion from a single final-round evaluation episode.
Figure 10: FedGuide-family qualitative rollouts on Hopper. Each cell overlays sampled frames from the final-round rollout to show the learned motion trace.
Figure 11: FedGuide-family qualitative rollouts on Walker2D, visualized with motion across heterogeneous clients.
Figure 12: FedGuide-family qualitative rollouts on HalfCheetah, visualized with motion overlays across the first eight clients.
Figure 13: FedGuide-family qualitative rollouts on MetaWorld10, visualized with across the ten task clients.
Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. Many current methods for PFRL rely heavily on exploiting existing reinforcement learning reward signals to derive an optimal policy for each client, thereby neglecting exploration in non-stationary or sparse-reward environments. In this work, we introduce a new exploration-driven framework, Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation (EDPFRL-IM), that leverages an inherent curiosity-driven exploration at each client to promote local exploration and protect client privacy. Furthermore, to facilitate policy discovery via exploration in previously unexplored state spaces, clients add an intrinsic random network distillation (RND) signal to their extrinsic reward. Additionally, the server does not have access to clients' raw experiences or local gradient estimates; instead, the server sends global exploration priors and collects minimal novelty summaries from each client to enable both diverse and coordinated exploration among clients. Experiments in benchmark environments show that our framework outperforms average PFRL benchmarks in policy personalization and sample efficiency, primarily in delayed and sparse reward systems. Overall, EDPFRL-IM enables the integration of a flexible exploratory learning structure into federated reinforcement learning systems while preserving client privacy.
Md Rafid Islam, Rafsan Jany, Zahid Hasan +1
Department of Electrical and Computer Engineering North South University Dhaka, Bangladesh · Digital Health Research Division Korea Institute of Oriental Medicine Daejeon, South Korea · Department of Electrical and Computer Engineering The University of Alabama in Huntsville Huntsville, AL, USA
This paper considers reinforcement learning from human feedback in a federated learning setting with resource-constrained agents, such as edge devices. We propose an efficient federated RLHF algorithm, named Partitioned, Sign-based Stochastic Zeroth-order Policy Optimization (Par-S2ZPO). The algorithm is built on zeroth-order optimization with binary perturbation, resulting in low communication, computation, and memory complexity by design. Our theoretical analysis establishes an upper bound on the convergence rate of Par-S2ZPO, revealing that it is as efficient as its centralized counterpart in terms of sample complexity but converges faster in terms of policy update iterations. Our experimental results show that it outperforms a FedAvg-based RLHF on four MuJoCo RL tasks.
Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories. However, FedAvg-style parameter averaging is not function-space consistent: when clients use heterogeneous encoders or even identical nonlinear networks, averaged parameters need not correspond to the weighted average of client value functions in any common function space. We propose FedQHD, a federated Q-learning method using hyperdimensional (random-feature) state encoders with a linear readout, so that Q-functions are nonlinear in state yet linear in trainable parameters. This linear structure enables closed-form aggregation. With a shared encoder, the function-space consensus update coincides exactly with weighted averaging of local readout matrices. With heterogeneous encoders, the server constructs a global teacher by averaging client Q-values on a shared anchor-state set, and each client compiles this teacher into its local representation via a single ridge projection. We formalize the federation gap -- the error incurred when compiling a federated teacher into a heterogeneous client representation -- relative to a client-specific oracle projection. We show that this gap decomposes into subspace misalignment, anchor-set conditioning, and regularization bias. We further identify the anchor-to-dimension ratio m≥Di as the well-conditioned regime in which the gap reduces to a multiple of the encoder heterogeneity floor. On four continuous-state, discrete-action control benchmarks, FedQHD matches or outperforms FedAvg-style baselines and distillation-based alternatives while requiring substantially less computation, and the empirical dependence of the federation gap on encoder dimension matches our theoretical analysis.
Yuchen Hou, Yongshan Chen, Zhuowen Zou +4
Northeastern University · University of California, Irvine · The George Washington University