Federated reinforcement learning (FRL) enables distributed agents to collaboratively train decision-making policies, but its decentralized training process also exposes global policy learning to Byzantine manipulation. Existing poisoning attacks primarily focus on how to construct malicious updates, while trajectory-level intervention timing remains largely implicit. In sequential decision making, however, where an intervention is applied can alter subsequent trajectories and learning signals. Through controlled experiments, we find that changing the selected trajectory states materially alters attack efficacy even when the malicious-update construction is fixed. We therefore identify when as a distinct attack dimension and introduce the Viability-constrained Behavioral Steering Attack (V-BSA), which uses local policy uncertainty to select sparse intervention states and applies envelope-constrained behavioral steering. Across discrete-action benchmarks, V-BSA achieves substantial degradation against robust aggregators and ensemble defenses with only a fraction of the interventions used by dense poisoning, while revealing task- and aggregation-dependent boundaries. Overall, our results highlight intervention timing as a distinct dimension of sequential robustness in FRL. The code is available at https://github.com/Yodeesy/V-BSA
Figures & tables
Figure 1: Conceptual comparison of intervention timing in FRL poisoning. (a) Prior work leaves trajectory-level intervention timing implicit, effectively defaulting to dense, state-agnostic intervention across rollout steps. (b) Ours ( V-BSA ) treats when as an explicit attack dimension: compromised agents maintain nominal execution on routine states and selectively intervene at states identified by policy-entropy gating.
Figure 2: Pilot evaluation of intervention timing. Terminal degradation D on LunarLander under FedPG-BR with fixed update construction across schedules ( ρ ). Entropy gating induces substantially higher degradation than dense and budget-matched baselines.
Figure 3: Overview of V-BSA . The attack factorizes into trajectory-level when selection and parameter-space how construction: an entropy gate selects sensitive states Di,whent , from which a targeted update is synthesized, norm-matched, and constrained by the local empirical envelope Bit before server-side aggregation.
Aggregation Rule
Ensemble ( K=5,M=10 )
Non-Ensemble ( N=30,M=10 )
LunarLander
Acrobot
CartPole
LunarLander
Acrobot
CartPole
FedAvg
79.90
120.50
3.00
−57.19
351.17
1.63
FedPG-BR
218.75
194.80
110.63
288.01
70.63
114.27
Coordinate-wise Median
175.83
151.17
8.80
434.35
284.87
42.80
Trimmed Mean
173.96
323.97
66.40
193.38
312.97
5.87
Table 1: Mean degradation D across aggregations and tasks (higher is more damaging). Positive values indicate successful terminal policy degradation induced by V-BSA .
Figure 4: Evaluation dynamics on Ensemble-LunarLander ( K=5,M=10 ). Returns of Clean vs. V-BSA across four aggregation rules, with shaded bands denoting standard deviation across random seeds.
Figure 5: Mechanistic audit of when selection (LunarLander, epoch 261). (a) Normalized policy entropy h(s) versus behavioral susceptibility L(s) . (b) Susceptibility distributions for non-gated versus gated states across five client groups.
Figure 6: Timing ablation under matched budgets (LunarLander × FedPG-BR). Fixing all how parameters, entropy gating achieves 1.89× higher mean degradation than budget-matched random selection across evaluated dynamics.
Figure 7: Temporal countermeasures against V-BSA on LunarLander. Terminal degradation D across dynamics d0,d1,d2 under no defense, the epoch auditor, and the temporal accumulator. The dashed line marks the preregistered resilience threshold ( 3σ=49.7 ).
Table 4: Discrete LunarLander-v2 evaluation breakdown under ensemble training ( K=5,M=10 ). Terminal evaluation returns and degradation D=Jctl−Jatk across evaluation dynamics.
Dynamics d0
Dynamics d1
Dynamics d2
Aggregation Rule
Jctl
Jatk
D
Jctl
Jatk
D
Jctl
Jatk
D
FedPG-BR
162.7
−59.2
221.91
240.4
−115.1
355.48
225.1
−61.5
286.64
Trimmed Mean
258.4
−12.3
270.70
249.1
−15.5
264.60
233.9
189.1
44.80
Median
234.4
−287.2
521.60
210.3
−168.1
378.40
225.8
−177.2
403.00
FedAvg
216.3
255.0
−38.70
198.2
269.1
−70.90
201.5
263.4
−61.90
Appendix
Table 5: Discrete LunarLander-v2 evaluation breakdown under non-ensemble training ( N=30,M=10 ). Terminal returns and degradation D across evaluation dynamics.
Table 6: Continuous control evaluations on LunarLanderContinuous-v2 under FedPG-BR. Terminal degradation D across evaluation dynamics under fixed and adaptive perturbation scaling.