Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especially in healthcare, where clinically consequential decisions require not only accuracy but also interpretable, auditable rationales to meet safety, accountability, and regulatory requirements. Conventional federated fine-tuning largely imitates final answers rather than cultivating step-by-step reasoning, often relying on privacy-sensitive centralized distillation and still incurring substantial communication overhead. We address this gap with \textbf{\ours{}}, a federated reasoning framework that combines lightweight chain-of-thought resampling with a compact discriminator for selection, and client-aware LoRA stacking with weighted classifier aggregation to accommodate heterogeneity while reducing aggregation noise and communication; clients generate candidate chains and supervision locally, and only lightweight modules are aggregated on the server. Experiments on medical reasoning benchmarks show consistent gains under tight resource budgets while keeping data local and respecting privacy, offering an interpretable and resource-efficient solution. Our code is made publicly available at https://github.com/DIaacKr/FedCoT
Figures & tables
Figure 1: Overview of FedCoT framework. Left: The data preparation for training the discriminator. Middle: The federated fine-tuning of discriminator without the participation of local LLM which is only used in data preparation. Right: The Optimal Discrimination at test time. “Down Proj.” corresponds to the A matrix in LoRA, and “Up Proj.” corresponds to the B matrix. “CLS" here denotes the classifier module of discriminator.
Datasets
Train
Test
Source
PubMedQA ( 2019 )
450
500
Experts
BioASQ ( 2015 )
494
124
Articles
MMLU ( 2020 )
1299
163
Examination
MedMCQA ( 2022 )
3000
4183
Examination
MedQA ( 2021 )
10178
1273
Examination
Table 1: The size and source of the medical QA datasets used in the experiment.
Method
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Avg.
Acc. (%)
Δ (%)
Acc. (%)
Δ (%)
Acc. (%)
Δ (%)
Acc. (%)
Δ (%)
Acc. (%)
Δ (%)
Acc. (%)
Δ (%)
LLaMA-3-8B-Instruct
37.90
—
29.80
—
27.20
—
38.70
—
9.20
—
28.56
—
+Self-Consistency
40.30
+2.40
31.50
+1.70
24.70
-2.50
41.10
+2.40
2.80
-6.40
28.08
-0.48
+FedIT
42.74
+4.84
47.29
+17.49
53.73
+26.53
71.17
+32.47
13.20
+4.00
45.63
+17.07
+FedFFA-LoRA
25.00
-12.90
39.47
+9.67
37.78
+10.58
57.67
+18.97
4.60
-4.60
32.90
+4.34
+FLORA
37.10
-0.80
42.05
+12.25
45.56
+18.36
61.96
+23.26
7.00
-2.20
38.73
+10.17
Table 2: Performance of different methods across five privacy-preserving medical datasets on top of two backbone LLMs under different settings. The best results are in Bold and the second-highest results are indicated with an underline .
Figure 2: Analysis of communication efficiency in federated SFT and our FedCoT. “SFT" represents FedIT, “Homo" represents FedCoT with lora rank of 32, “Heter" represents FedCoT with lora rank of 4, 32, 32, 16, 4.
Method
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Avg.
Best-of-N (Logits-based)
40.32
27.76
26.24
38.04
5.60
27.59
Best-of-N (Qwen3-0.6B-Reranker)
49.20
26.50
29.60
36.20
12.40
30.78
FedCoT (Local)
57.30
42.60
40.80
56.40
29.20
45.26
FedCoT (Qwen3-0.6B-Base)
50.00
43.70
37.00
49.70
13.20
38.72
FedCoT (ours)
68.50
45.20
54.10
68.70
41.00
55.50
Table 3: Performance of different filtering method on 8 candidates. FedCoT (ours) denotes the full federated framework. FedCoT (Local) represents models trained only on individual client datasets. FedCoT (Qwen3-0.6B-Base) replaces the default discriminator with a generative backbone. All results are accuracy (%).
Method
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Avg.(%)
CoT
37.90
29.80
27.20
38.70
9.20
28.56
FedCoT (r=4,4,4,4,4)
67.70
44.50
52.60
65.00
40.80
54.12
FedCoT (r=8,8,8,8,8)
68.50
45.40
53.40
65.60
41.00
54.78
FedCoT (r=16,16,16,16,16)
69.40
45.00
54.30
65.60
41.00
55.06
FedCoT (r=32,32,32,32,32)
66.10
45.70
54.00
66.30
41.00
54.62
FedCoT (r=4,32,32,16,4)
68.50
45.20
54.10
68.70
41.00
55.50
Table 4: The different performances of FedCoT under different LoRA configurations. “r" represents the LoRA rank of different clients, corresponding to the clients of the datasets BioASQ, MedMCQA, MedQA, MMLU, and PubMedQA in sequence. The best results are in Bold .
Method
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Avg.
Acc. (%)
Size
Acc. (%)
Size
Acc. (%)
Size
Acc. (%)
Size
Acc. (%)
Size
Acc. (%)
FedCoT (Full)
68.50
364
45.20
1420
54.10
7278
68.70
915
41.00
205
55.50
w/o Candidate Selection
69.40
494
44.90
2000
53.70
10178
62.90
1299
41.00
450
54.38
w/o Soft Labeling
67.70
364
45.10
1420
54.50
7278
60.30
915
41.00
205
53.72
w/o CS & SL
65.30
494
45.20
2000
56.10
10178
54.00
1299
41.00
450
52.32
Table 5: Ablation study of FedCoT components. “Size” denotes the number of candidate paths retained for local discriminator training.
Figure 3: Impact of varying the MMLU client’s LoRA rank on local accuracy, communication cost, and the average performance of other fixed-rank clients ( r=8 ).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Max Length
Actor Model
Avg.(%)
512
LLaMA-3-8B-Instruct
5.93
Qwen2.5-7B-Instruct
26.28
1024
LLaMA-3-8B-Instruct
0.02
Qwen2.5-7B-Instruct
0.04
Appendix
Table 6: The mean truncation rates on all datasets across maximum generation lengths.
Model
Method
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Avg.(%)
LLaMA-3-8B- Instruct
FedIT-512
42.74
47.29
53.73
71.17
13.20
45.63
FedCoT-512
65.30
45.20
56.10
54.00
41.00
52.32
FedIT-1024
45.16
46.45
54.36
70.55
13.80
46.06
FedCoT-1024
78.20
44.40
53.70
58.30
42.20
55.36
Qwen2.5-7B- Instruct
FedIT-512
82.26
48.48
44.30
68.71
47.20
58.19
FedCoT-512
96.80
50.00
52.50
66.30
64.80
66.08
Appendix
Table 7: Performance comparison (%) across different maximum generation token lengths. The best results are in Bold and the second-highest results are indicated with an underline .
Dataset
Positive
Negative
Ratio(%)
BioASQ
5,740
145
97.54
MedMcQA
47,290
5,408
89.74
MedQA
118,719
26
99.98
MMLU
36,676
979
97.40
PubMedQA
3,288
740
81.63
Appendix
Table 8: Step-wise self-evaluation performance across medical QA benchmarks. Positive: count of reasoning steps judged correct by the model; Negative: count of steps judged incorrect; Ratio: proportion of correct self-evaluation.
Model
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Baseline
37.90
29.80
27.20
38.70
9.20
Skywork-Reward-V2-Qwen3-0.6B
31.50
24.50
20.20
32.50
4.00
Skywork-Reward-V2-Llama-3.2-1B
41.90
23.90
25.80
28.80
6.40
Skywork-Reward-V2-Qwen3-1.7B
25.00
24.80
21.20
26.40
1.00
Skywork-Reward-V2-Llama-3.1-8B
16.10
12.80
7.60
16.00
0.60
Appendix
Table 9: Performance of open-source SOTA reward models in the medical domain across different benchmarks (in %).
Question :A previously healthy 29-year-old man comes to the emergency department because of burning with urination for several days. He has also had pain in the right ankle for 3 days and pain and swelling in the left knee for 1 day. Two weeks ago, he had several days of fever and bloody diarrhea, for which he was treated with antibiotics. Examination shows a small left knee effusion and bilateral conjunctival injection. Which of the following is the most likely additional finding in this patient?\n A. Palpable mass in the right lower quadrant\n B. Tenderness at the insertion of the Achilles tendon\n C. Pain on passive extension of the fingers\n D. Circular erythematous rash with central clearing Ground-Truth Answer : B ;
Generation 1 : Here is the step-by-step reasoning to answer the question:\n\nThe patient presents with a combination of symptoms, including burning with urination, pain in the right ankle, and pain and swelling in the left knee. … The development of new symptoms, such as joint pain and swelling, suggests that the infection may have spread to other parts of the body.\n\nThe patient’s conjunctival injection (inflammation of the conjunctiva) is also a significant finding, as it is a common manifestation of disseminated infection.\n\nGiven the patient’s symptoms and medical history, the most likely diagnosis is reactive arthritis, which is a type of spondyloarthropathy that occurs in response to an underlying infection, often gastrointestinal.\n\nReactive arthritis typically affects the joints, eyes, and skin, and can cause symptoms such as joint pain, swelling, and inflammation. The patient’s bilateral conjunctival injection and left knee effusion are consistent with this diagnosis.\n\nThe most likely additional finding in this patient would be:\n\n#### C. Pain on passive extension of the fingers\n\nThis symptom is a common manifestation of reactive arthritis, as the condition can cause inflammation and stiffness in the hands and fingers. Passive extension of the fingers would likely cause pain and stiffness, which would be consistent with the patient’s underlying diagnosis.\n\nThe other options are not as likely:\n\nB. Tenderness at the insertion of the Achilles tendon: Achilles tendonitis is a possible diagnosis, but it is not the most likely explanation for the patient’s symptoms given the patient’s medical history and conjunctival injection.\n\n
Prediction 1 : C ; Score 1: 0.77344
Generation 2 : Here is the step-by-step reasoning to answer the question:\n\nThe patient presents with a complex set of symptoms, including burning with urination, pain in the right ankle, and pain and swelling in the left knee. The patient also reports a history of fever and bloody diarrhea two weeks ago, which was treated with antibiotics.\n\n … Step 4: Consider the possible causes of the patient’s symptoms in the context of the possible sources of infection.\n\nGiven the patient’s history of bloody diarrhea and the presence of symptoms in the joints and eyes, it is possible that the patient has developed a septic arthritis, which is an infection of the joint that is caused by bacteria that have spread from another part of the body.\n\nStep 5: Evaluate the patient’s symptoms in relation to the possible causes of septic arthritis.\n\nThe patient’s symptoms of pain and swelling in the joints, as well as the presence of a small left knee effusion, are consistent with septic arthritis. The patient’s bilateral conjunctival injection also suggests that the infection may have spread to the eyes.\n\nStep 6: Consider the most likely additional finding in the patient.\n\nGiven the patient’s symptoms and the possible causes of septic arthritis, the most likely additional finding is a palpable mass in the right lower quadrant, which could indicate an abscess or a collection of pus in the right lower quadrant of the abdomen.\n\n#### A. Palpable mass in the right lower quadrant
Prediction 2 : A ; Score 2: 0.717285
Generation 3 : Here is the step-by-step reasoning:\n\nThe patient presents with a combination of symptoms, including burning with urination, pain in the right ankle, and pain and swelling in the left knee. This suggests that the patient may have a systemic infection that is affecting multiple parts of the body.\n\nThe patient had a recent episode of bloody diarrhea, which is a common symptom of gastrointestinal infections. The fact that he was treated with antibiotics suggests that the infection was likely bacterial in nature.\n\nThe patient’s symptoms are consistent with a urinary tract infection (UTI), as burning with urination is a common symptom of UTIs. The presence of a small left knee effusion and bilateral conjunctival injection suggests that the infection may have spread to the joints and eyes, respectively.\n\nThe patient’s symptoms are also consistent with reactive arthritis, a condition that can occur as a complication of certain bacterial infections, such as Salmonella or Shigella. Reactive arthritis is characterized by joint pain and swelling, as well as conjunctivitis and urethritis.\n\nGiven the patient’s symptoms and medical history, the most likely additional finding is:\n\nB. Tenderness at the insertion of the Achilles tendon\n\nThis is because reactive arthritis often affects the joints, including the ankles, and the Achilles tendon is a common site of tenderness in this condition.\n\n#### B. Tenderness at the insertion of the Achilles tendon
Prediction 3 : B ; Score 3: 0.833984
Appendix
Table 10: A case study on MedQA dataset.
Model
bioasq
medmcqa
medqa
mmlu
pubmedqa
average
Δ
p@1
p@k
p@1
p@k
p@1
p@k
p@1
p@k
p@1
p@k
p@1
p@k
LLaMA-3-8B
37.90
70.97
29.84
68.18
27.18
70.46
38.65
71.78
9.20
42.00
28.55
64.68
36.12
LLaMA-3.1-8B
25.81
64.52
35.07
72.22
32.05
73.53
39.26
77.30
15.20
54.80
29.48
68.47
38.99
Qwen2.5-7B
73.39
98.39
43.75
70.98
29.46
66.54
50.31
78.53
38.80
73.40
47.14
77.57
30.43
Qwen2.5-3B
10.48
46.77
18.98
55.41
5.03
23.80
26.38
65.64
1.20
7.40
12.41
39.81
27.39
LLaMA3.2-3B
31.45
70.16
31.41
65.57
20.03
59.62
32.52
70.55
4.80
24.20
24.04
58.02
33.98
Appendix
Table 11: Model performance (in %) on biomedical test sets demonstrating accuracy potential through multiple sampling. The table shows pass rates at first sample (p@1) and after k samples (p@k), with Δ=p@k−p@1 indicating accuracy improvement potential.
Method
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Avg.
LlaMA-3-8B-Instruct
37.90
29.80
27.20
38.70
9.20
28.56
FedCoT (5 clients)
68.50
45.20
54.10
68.70
41.00
55.50
FedCoT (8 clients)
69.40
44.30
52.80
69.30
41.20
55.40
Appendix
Table 12: Performance under different client granularities. The eight-client setting partitions MedMCQA into two clients and MedQA into three clients. All results are accuracy (%).
Method
BioASQ
MedMCQA
MedQA
MMLU
PubMedQA
Avg.
Qwen2.5-7B-Instruct
73.40
43.70
29.50
50.30
38.80
47.14
FedCoT (shared head)
96.00
51.70
53.40
74.20
68.20
68.70
FedCoT (personalized head)
80.60
48.20
29.50
60.10
43.60
52.40
Appendix
Table 13: Effect of classifier-head aggregation. The local-head variant keeps the classifier head client-specific and aggregates only LoRA modules. All results are accuracy (%).
Figure 6: Performance improvement on difference candidate numbers of FedCoT.
Figure 7: LLM-as-Judge evaluation on four biomedical QA datasets across three axes (higher is better): medical factuality, explainability/coherence, and overall quality.
Figure 8: Performance improvement on 3B LLMs via federated reasoning fine-tuning on top of our FedCoT.
Figure 9: The validation in out-of-distribution situation.
Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across organizations that cannot centralize it due to regulatory, proprietary, or institutional constraints. We study federated reasoning, where a server improves multi-step reasoning by coordinating with heterogeneous clients holding private demonstrations, without centralized training or raw data sharing. The key challenge is that client reliability is query-dependent, while the server cannot inspect client data to determine which contributions are trustworthy. To address this, we propose Uncertainty-Aware Federated Reasoning (FERA), a training-free framework based on iterative server-client co-refinement. Across communication rounds, clients generate reasoning traces with lightweight uncertainty estimates, and the server synthesizes them into improved reasoning that is redistributed as context for the next round, progressively improving both server outputs and client-side reasoning. Within each round, Uncertainty-Aware Self-Critique Aggregation (UA-SCA) resolves conflicts among heterogeneous client traces through query-dependent trust weighting and structured cross-client verification. Rather than simply discarding low-quality traces, UA-SCA revises flawed reasoning steps to recover useful information. We provide theoretical guarantees showing that the proposed iterative protocol converges and that uncertainty-aware weighting accelerates convergence. Experiments on multiple reasoning benchmarks show that FERA consistently outperforms both federated training and training-free baselines, achieving progressively higher accuracy across rounds while maintaining communication and computational efficiency.
Ruhan Wang, Chengkai Huang, Zhiyong Wang +6
Indiana University · The University of New South Wales · The Chinese University of Hong Kong +2
Modern agents specialize in varying domains while there is no clear approach combining different domain skills. We propose a federated learning-like framework, Federation over Text (FoT), that enables multiple clients solving different tasks to collectively generate a shared library of metacognitive insights by iteratively federating their local reasoning processes without sharing actual problem instances. Instead of federation over gradients (e.g., as in distributed training), FoT operates at the semantic level without any gradient optimization or supervision signal. At each round, client LLM agents independently apply arbitrary local reasoning and self-improvement procedures to their own tasks and share the resulting reasoning traces with a central server. The server then aggregates, distills, and consolidates knowledge across tasks and domains into a shared insight library, which can be reused by current and future agents to improve their reasoning. Experiments show that FoT improves reasoning effectiveness and efficiency across real-world daily tasks, cross-domain collaboration, and research insight discovery, achieving an average performance gain of 11.9 absolute percentage points while reducing client-generated task-completion tokens by 5.5%.
Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods require transmitting large collections of trainable weights, assume aligned architectures, and rely on white-box access to model parameters. As model sizes continue to grow and deployments become increasingly heterogeneous, these assumptions become progressively misaligned with practical constraints. We consider an alternative formulation in which collaboration is mediated through model behavior rather than parameters. Clients fine-tune local models on private data and exchange generated outputs on a shared, public prompt set. The server maps these outputs into a semantic representation space, forms a per-prompt semantic consensus, and returns pseudo-labels for further local fine-tuning. This formulation fundamentally changes the communication scaling of federated LLM fine-tuning. The amount of information exchanged depends only on the public prompt budget and the size of the communicated behaviors, independent of model size. As a consequence, the protocol naturally accommodates heterogeneous architectures and applies directly to open-ended text generation. We present a theoretical analysis and empirical results demonstrating that this approach can match strong federated fine-tuning baselines while substantially reducing communication by orders of magnitude (e.g., analytically by a factor of 1006 for Llama3.1-405B), as well as reductions in runtime and energy consumption. These results suggest that, for generative foundation models, behavior-level consensus provides a more appropriate abstraction for federated adaptation than parameter aggregation.
Amr Abourayya, Jens Kleesiek, Michael Kamp
Lamarr Institute for ML and AI, Technical University Dortmund · Institute for AI in medicine (IKIM),University Hospital Essen