Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation, prioritizing clients with stronger learning signals; (ii) \emph{global reward calibration} that re-weights per-prompt objectives based on the local-global reward gap, steering each client toward its relative weaknesses; and (iii) \emph{adaptive sparse communication} that allocates bandwidth based on the informativeness of each client's update. Extensive experiments on mathematical reasoning tasks demonstrate that Fed-GRPO achieves the best performance among all federated methods, clearly outperforms FedAvg and approaches centralized training performance, while losslessly reducing communication by 32× and supporting up to 621× compression under tight bandwidth budgets with only graceful accuracy degradation. Our code is available at https://github.com/HKU-HealthAI/Fed-GRPO.
Figures & tables
Figure 1: Motivation of Fed-GRPO. Naïve FedAvg-GRPO suffers from three coupled issues: (1) reward-misaligned aggregation, (2) local–global objective miscalibration, and (3) reward-agnostic communication. We observe that the reward statistics (μk,σk) , computed by GRPO at zero extra cost , are precisely the signals needed to address all three: σk guides both signal-weighted aggregation and adaptive bandwidth allocation, while the local–global gap on μk calibrates per-prompt objectives.
Figure 2: FedAvg-GRPO training dynamics on the MATH dataset with 5 clients split by difficulty. (a) Per-client reward standard deviation σk varies by up to 4× across clients within the same round, indicating large differences in learning signal strength. (b) Per-client reward mean μk diverges significantly across clients, revealing a persistent local-global gap. (c) Average model update sparsity across clients: over 96.2% of parameters remain unchanged after local GRPO training. (d) Correlation between σk and model update sparsity: clients with stronger learning signals produce denser updates.
Method
Budget
Comm. (MB)
Comp. ( × )
GSM8K
MATH-500
AIME24
AIME25
AMC
Avg.
Qwen2.5-3B-Instruct
-
-
-
59.47
46.54
2.50
1.46
27.11
27.42
Centralized
-
-
-
84.75
63.42
7.50
3.44
41.95
40.21
Local-Only
-
-
-
82.45
58.16
5.88
2.94
38.28
37.54
FedAvg-SW
-
6,794.2
1×/1×
79.55
56.84
6.56
3.02
39.53
37.10
FedAvg-UW
-
6,794.2
1×/1×
80.00
55.73
5.42
3.12
37.66
36.39
SparsyFed
5%†
1,019.1
20×/7×
79.10
56.61
6.56
2.92
39.53
36.94
Table 1: Main results on five math reasoning benchmarks using Qwen2.5-3B-Instruct as the base model . FedAvg-SW and FedAvg-UW denote sample-weighted and uniform-weighted FedAvg, respectively. Budget B is the target fraction of changed parameters retained for transmission; Comm. (MB) is the average per-round per-client uplink payload (BF16 values plus 4-byte indices for sparse encodings); Comp. ( × ) reports the compression ratio relative to dense full-model transmission as parameter ratio / byte ratio . Best result for federated methods in bold . † For SparsyFed and SparseLoCo, Budget denotes the kept fraction of all parameters.
Table 4
Configuration
GSM8K
MATH-500
AIME24
AIME25
AMC
Avg.
Comm. (MB)
FedAvg-SW (no component)
79.55
56.84
6.56
3.02
39.53
37.10
6,794.2
+ SWA (§ 4.1 )
81.24
58.73
7.02
3.05
40.03
38.01
6,794.2
+ SWA + GRC (§ 4.2 )
82.75
61.26
7.26
3.13
40.94
39.07
6,794.2
+ SWA + GRC + ASC (Fed-GRPO)
83.21
60.82
7.08
3.12
41.09
39.06
645.5
Table 4: Ablation studies on Fed-GRPO components. + SWA: add signal-weighted aggregation. + GRC: add global reward calibration. + ASC: add adaptive sparse communication (at B=1.0 , lossless). Comm. (MB) refers to the average per-round per-client uplink payload.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
GSM8K
MATH-500
AIME24
AIME25
AMC
Avg.
Qwen2.5-3B-Instruct
59.47
46.54
2.50
1.46
27.11
27.42
Centralized
84.75
63.42
7.50
3.44
41.95
40.21
Local-Only
79.60
53.09
4.43
2.26
33.57
34.59
FedAvg-SW
81.32
56.04
5.04
2.72
35.12
36.05
FedAvg-UW
80.79
56.23
4.97
2.66
34.86
35.90
SparsyFed
80.96
56.02
5.02
2.53
35.12
35.93
Appendix
Table 5: Results on five math reasoning benchmarks using Qwen2.5-3B-Instruct as the base model under dirichlet partitioning ( α =0.5) by ”subject” with K=20 clients.
Figure 3: Dynamics of (a) per-client reward standard deviation σk and (b) aggregation weights wk . The aggregation weights shift dynamically across rounds, concentrating on clients with stronger learning signals (higher σk ) as training progresses.
Recent advances in language models have established reinforcement learning as the primary paradigm for eliciting self-correction and long-chain reasoning. While group relative policy optimization (GRPO) offers superior scalability by eliminating the critic network, deploying it on a central infrastructure entails collecting a large volume of data from distributed owners, which poses significant privacy risks. To address these concerns, we introduce federated GRPO (FGRPO), a framework designed to decentralize the fine-tuning of reasoning models across heterogeneous data owners. To effectively mitigate the instability caused by divergent reward scales across heterogeneous tasks, FGRPO incorporates an adaptive aggregation mechanism based on relative performance gain. By characterizing each client's improvement relative to its personalized historical baseline, the framework dynamically prioritizes effective learning trajectories regardless of local task difficulty. FGRPO ensures robust convergence on non-IID data while preserving data privacy.
Pengyu Chen, Shaowei Li, Kai Wang +4
School of Computer Science and Technology, Shandong University, Qingdao China · School of Mathematical Science, Peking University, China · School of Computer Science and Artificial Intelligence, Shanghai University of Finance and Economics, Shanghai, China +1
Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especially in healthcare, where clinically consequential decisions require not only accuracy but also interpretable, auditable rationales to meet safety, accountability, and regulatory requirements. Conventional federated fine-tuning largely imitates final answers rather than cultivating step-by-step reasoning, often relying on privacy-sensitive centralized distillation and still incurring substantial communication overhead. We address this gap with \textbf{\ours{}}, a federated reasoning framework that combines lightweight chain-of-thought resampling with a compact discriminator for selection, and client-aware LoRA stacking with weighted classifier aggregation to accommodate heterogeneity while reducing aggregation noise and communication; clients generate candidate chains and supervision locally, and only lightweight modules are aggregated on the server. Experiments on medical reasoning benchmarks show consistent gains under tight resource budgets while keeping data local and respecting privacy, offering an interpretable and resource-efficient solution. Our code is made publicly available at https://github.com/DIaacKr/FedCoT
Chuan Li, Qianyi Zhao, Fengran Mo +1
East China Normal University, China · University of Montreal, Canada
Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across organizations that cannot centralize it due to regulatory, proprietary, or institutional constraints. We study federated reasoning, where a server improves multi-step reasoning by coordinating with heterogeneous clients holding private demonstrations, without centralized training or raw data sharing. The key challenge is that client reliability is query-dependent, while the server cannot inspect client data to determine which contributions are trustworthy. To address this, we propose Uncertainty-Aware Federated Reasoning (FERA), a training-free framework based on iterative server-client co-refinement. Across communication rounds, clients generate reasoning traces with lightweight uncertainty estimates, and the server synthesizes them into improved reasoning that is redistributed as context for the next round, progressively improving both server outputs and client-side reasoning. Within each round, Uncertainty-Aware Self-Critique Aggregation (UA-SCA) resolves conflicts among heterogeneous client traces through query-dependent trust weighting and structured cross-client verification. Rather than simply discarding low-quality traces, UA-SCA revises flawed reasoning steps to recover useful information. We provide theoretical guarantees showing that the proposed iterative protocol converges and that uncertainty-aware weighting accelerates convergence. Experiments on multiple reasoning benchmarks show that FERA consistently outperforms both federated training and training-free baselines, achieving progressively higher accuracy across rounds while maintaining communication and computational efficiency.
Ruhan Wang, Chengkai Huang, Zhiyong Wang +6
Indiana University · The University of New South Wales · The Chinese University of Hong Kong +2