Federated reinforcement learning (FRL) enables distributed agents to collaboratively train decision-making policies, but its decentralized training process also exposes global policy learning to Byzantine manipulation. Existing poisoning attacks primarily focus on how to construct malicious updates, while trajectory-level intervention timing remains largely implicit. In sequential decision making, however, where an intervention is applied can alter subsequent trajectories and learning signals. Through controlled experiments, we find that changing the selected trajectory states materially alters attack efficacy even when the malicious-update construction is fixed. We therefore identify when as a distinct attack dimension and introduce the Viability-constrained Behavioral Steering Attack (V-BSA), which uses local policy uncertainty to select sparse intervention states and applies envelope-constrained behavioral steering. Across discrete-action benchmarks, V-BSA achieves substantial degradation against robust aggregators and ensemble defenses with only a fraction of the interventions used by dense poisoning, while revealing task- and aggregation-dependent boundaries. Overall, our results highlight intervention timing as a distinct dimension of sequential robustness in FRL. The code is available at https://github.com/Yodeesy/V-BSA
Figures & tables
Figure 1: Conceptual comparison of intervention timing in FRL poisoning. (a) Prior work leaves trajectory-level intervention timing implicit, effectively defaulting to dense, state-agnostic intervention across rollout steps. (b) Ours ( V-BSA ) treats when as an explicit attack dimension: compromised agents maintain nominal execution on routine states and selectively intervene at states identified by policy-entropy gating.
Figure 2: Pilot evaluation of intervention timing. Terminal degradation D on LunarLander under FedPG-BR with fixed update construction across schedules ( ρ ). Entropy gating induces substantially higher degradation than dense and budget-matched baselines.
Figure 3: Overview of V-BSA . The attack factorizes into trajectory-level when selection and parameter-space how construction: an entropy gate selects sensitive states Di,whent , from which a targeted update is synthesized, norm-matched, and constrained by the local empirical envelope Bit before server-side aggregation.
Aggregation Rule
Ensemble ( K=5,M=10 )
Non-Ensemble ( N=30,M=10 )
LunarLander
Acrobot
CartPole
LunarLander
Acrobot
CartPole
FedAvg
79.90
120.50
3.00
−57.19
351.17
1.63
FedPG-BR
218.75
194.80
110.63
288.01
70.63
114.27
Coordinate-wise Median
175.83
151.17
8.80
434.35
284.87
42.80
Trimmed Mean
173.96
323.97
66.40
193.38
312.97
5.87
Table 1: Mean degradation D across aggregations and tasks (higher is more damaging). Positive values indicate successful terminal policy degradation induced by V-BSA .
Figure 4: Evaluation dynamics on Ensemble-LunarLander ( K=5,M=10 ). Returns of Clean vs. V-BSA across four aggregation rules, with shaded bands denoting standard deviation across random seeds.
Figure 5: Mechanistic audit of when selection (LunarLander, epoch 261). (a) Normalized policy entropy h(s) versus behavioral susceptibility L(s) . (b) Susceptibility distributions for non-gated versus gated states across five client groups.
Figure 6: Timing ablation under matched budgets (LunarLander × FedPG-BR). Fixing all how parameters, entropy gating achieves 1.89× higher mean degradation than budget-matched random selection across evaluated dynamics.
Figure 7: Temporal countermeasures against V-BSA on LunarLander. Terminal degradation D across dynamics d0,d1,d2 under no defense, the epoch auditor, and the temporal accumulator. The dashed line marks the preregistered resilience threshold ( 3σ=49.7 ).
Table 4: Discrete LunarLander-v2 evaluation breakdown under ensemble training ( K=5,M=10 ). Terminal evaluation returns and degradation D=Jctl−Jatk across evaluation dynamics.
Dynamics d0
Dynamics d1
Dynamics d2
Aggregation Rule
Jctl
Jatk
D
Jctl
Jatk
D
Jctl
Jatk
D
FedPG-BR
162.7
−59.2
221.91
240.4
−115.1
355.48
225.1
−61.5
286.64
Trimmed Mean
258.4
−12.3
270.70
249.1
−15.5
264.60
233.9
189.1
44.80
Median
234.4
−287.2
521.60
210.3
−168.1
378.40
225.8
−177.2
403.00
FedAvg
216.3
255.0
−38.70
198.2
269.1
−70.90
201.5
263.4
−61.90
Appendix
Table 5: Discrete LunarLander-v2 evaluation breakdown under non-ensemble training ( N=30,M=10 ). Terminal returns and degradation D across evaluation dynamics.
Table 6: Continuous control evaluations on LunarLanderContinuous-v2 under FedPG-BR. Terminal degradation D across evaluation dynamics under fixed and adaptive perturbation scaling.
Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard aggregation methods such as FedAvg do not account for system-level constraints, often leading to unsafe global behavior. In this work, we study constraint-aware aggregation for federated reinforcement learning in distributed energy coordination. We propose aggregation rules that incorporate both local performance and estimated constraint violation into the server-side update. Among these, a simple penalty-based rule, wi∝Ri−αVi, consistently provides the most reliable trade-off between reward and safety, without requiring dual optimization or modifications to local training. \textcolor{black}{We evaluate our approach on DairyGridEnv, a benchmark modeling multiple farms coordinating battery storage under stochastic demand and a shared grid capacity constraint, and further assess robustness using real load-driven demand profiles from Finland and the German FIELD dataset. Across multiple seeds, penalty-based aggregation substantially reduces violations while improving reward relative to FedAvg in both synthetic and real load-driven settings.} A combined reward-violation scheme exposes a tunable trade-off via λ, but is less stable. These results demonstrate that lightweight aggregation strategies can substantially improve empirical safety in federated reinforcement learning while preserving standard communication protocols.
Usman Haider, Karl Mason
School of Computer Science, University of Galway, Ireland.
This paper considers reinforcement learning from human feedback in a federated learning setting with resource-constrained agents, such as edge devices. We propose an efficient federated RLHF algorithm, named Partitioned, Sign-based Stochastic Zeroth-order Policy Optimization (Par-S2ZPO). The algorithm is built on zeroth-order optimization with binary perturbation, resulting in low communication, computation, and memory complexity by design. Our theoretical analysis establishes an upper bound on the convergence rate of Par-S2ZPO, revealing that it is as efficient as its centralized counterpart in terms of sample complexity but converges faster in terms of policy update iterations. Our experimental results show that it outperforms a FedAvg-based RLHF on four MuJoCo RL tasks.
While federated learning enables collaborative modelling on decentralised data, standard methods merely fit historical observations. This purely observational approach is fundamentally insufficient for interventional inference and policy evaluation, as sequential actions dynamically alter future states. We propose \textbf{Fed-CausalDiff}, a federated causal diffusion framework for do-simulation. The architecture decomposes the evolution of the latent state into a global causal score function and a local confounding score function. This design enables \emph{decoupled synchronisation} (DSS), where clients aggregate only the shared causal mechanism while retaining site-specific confounders locally to handle heterogeneity. Experiments on four datasets demonstrate that Fed-CausalDiff achieves better ATE and policy-value estimation accuracy, offering a favorable trade-off between communication cost and inference fidelity.
Pengfei Li, Mohammad Khalil
Centre for the Science of Learning & Technology (SLATE), University of Bergen Bergen, Norway