cs.AIOct 8, 2026

Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization

Authors: Pengxin Guo, Shuang Zeng, Zonggen Li, Weiying Zheng, Mengting Liu, Liangqiong Qu

Organizations: The University of Hong Kong · Sun Yat-sen University

Abstract

Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation, prioritizing clients with stronger learning signals; (ii) \emph{global reward calibration} that re-weights per-prompt objectives based on the local-global reward gap, steering each client toward its relative weaknesses; and (iii) \emph{adaptive sparse communication} that allocates bandwidth based on the informativeness of each client's update. Extensive experiments on mathematical reasoning tasks demonstrate that Fed-GRPO achieves the best performance among all federated methods, clearly outperforms FedAvg and approaches centralized training performance, while losslessly reducing communication by 32×32\times and supporting up to 621×621\times compression under tight bandwidth budgets with only graceful accuracy degradation. Our code is available at https://github.com/HKU-HealthAI/Fed-GRPO.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data

    Jun 2, 2026Pengyu Chen, Shaowei Li, Kai Wang +4Federated Learning AggregationFederated RL

  2. FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models

    Aug 7, 2025Chuan Li, Qianyi Zhao, Fengran Mo +1Communication-Efficient Distributed TrainingMedical LLMs

  3. FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

    May 11, 2026Ruhan Wang, Chengkai Huang, Zhiyong Wang +6LLM ReasoningFederated Learning