This work presents SeedFlood, a new approach to decentralized LLM fine-tuning designed to scale across large models, large collaborations, and complex network topologies while achieving global consensus with negligible communication overhead. Traditional methods suffer from high communication costs that grow with model size, while information decay over network hops renders global consensus inefficient. SeedFlood takes a significant departure from these practices by exploiting the seed-reconstructible structure of zeroth-order gradients and effectively making the messages to transmit near-zero in size, allowing them to be flooded to every client in the network, and thereby enhancing scalability of decentralized training. Consequently, SeedFlood enables training in regimes previously considered impractical, such as billion-parameter scale models or distributed across hundred of clients. Our experiments on decentralized LLM fine-tuning demonstrate that SeedFlood consistently outperforms the standard zeroth-order baselines in both communication efficiency and generalization performance, and even achieves results comparable to first-order gossip-based methods in large-scale settings, while requiring orders-of-magnitude less communication cost. We also provide theoretical analysis to formalize that SeedFlood avoids topology-dependent consensus terms in the convergence bound while retaining the acceleration enabled by increased client participation.
Figures & tables
Figure 1: Comparison of gossip averaging and S eed F lood. (a) Gossip reaches consensus through repeated neighbor-wise model averaging. (b) S eed F lood instead floods every client’s gradient to the entire network, so that each client updates with the average of all of them. (c) Flooding remains duplicate-free because each client forwards a message exactly once, upon its first reception. (d) Each zeroth-order gradient can be communicated using only a random seed and a scalar coefficient. (e) This lossless flooding enables rapid multi-hop propagation and exact consensus without the information dilution inherent in gossip averaging.
Figure 2: Subspace coordinate-wise perturbation of S eed F lood. (a) Instead of drawing perturbations from Gaussian distribution, S eed F lood prepares shared bases Ul,Vl , samples a random coordinate pair (al,bl) at each iteration, which selects the dense rank-one direction Uℓ,:,aℓVℓ,:,bℓ⊤ as the perturbation. (b) The n received scalars are accumulated into Aℓ and applied by a single matrix product.
Figure 3: Performance–communication trade-off on a 16-client ring network. (a,b) S eed F lood versus 9 decentralized fine-tuning baselines on OPT-1.3B and Qwen3-1.7B-Base; performance is averaged over tasks relative to DSGD-FFT. (c) Per-task performance of S eed F lood and representative baselines at the communication-cost extremes: DSGD-FFT, DZSGD-FFT, and ChocoSGD-FFALoRA. S eed F lood outperforms zeroth-order and the most communication-efficient first-order PEFT baselines with model-size-independent communication and 10 – 107× fewer communicated bytes.
Figure 4: Performance scaling with increasing network size in a ring topology. Unlike first-order gossip baselines, S eed F lood maintains stable or improving performance as the number of clients increases, and surpasses the best-performing gossip baseline at larger client counts.
Figure 5: Iterations required to reach the target loss across 16–128 clients, measured on training (a, c) and validation (b, d) sets. While S eed F lood initially requires more iterations than convergent first-order gossip baselines, its iteration count decreases with increasing client participation. In contrast, first-order gossip baselines require more iterations or fail to reach the target within the 10,000-iteration budget (dashed horizontal line).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Approach
Communicated Bytes
Applying Computation
Perfect Consensus
Traditional Gossip
O(d)
O(d)
–
S eed F lood (ours)
O(n)
O(n+rd)
✓
Appendix
Table 1: Communication overhead and consensus properties.
Figure 6: Performance–communication trade-off on a 16-client meshgrid network.
Figure 8
Figure 9: Task performances on SST2 and RTE under partial flooding with varying number of flooding steps.
Method
IID Acc.
non-IID Acc.
Total Comm.
Comm. / Iter.
DSGD
97.57
96.08
260 MB
173 KB
ChocoSGD
90.47
69.01
5.24 MB
3.5 KB
S eed F lood
94.17
92.82
2.78 MB
0.3 KB
DZSGD
94.10
84.45
1.3 GB
173 KB
Appendix
Table 3: Training from scratch on MNIST with a 3-layer CNN under a 32-client ring topology. S eed F lood achieves competitive accuracy while requiring substantially less communication than dense decentralized baselines.
IFEval
Comm.
Prompt
Inst.
Prompt
Inst.
Method
(MB total)
strict
strict
loose
loose
BBH
MMLU
Qwen3-8B-Base (untuned)
–
46.4
58.5
52.1
63.2
47.4
69.1
DSGD-FFALoRA
182.9
51.1 ± 0.3
61.5 ± 0.5
55.3 ± 0.3
65.3 ± 0.5
49.0 ± 0.1
71.5 ± 0.2
ChocoSGD-FFALoRA
5.5
50.3 ± 1.7
61.0 ± 1.6
55.0 ± 1.2
64.9 ± 1.2
48.9 ± 0.3
71.7 ± 0.2
DZO-FFALoRA
182.9
46.7 ± 0.7
58.4 ± 0.2
52.1 ± 0.7
62.7 ± 0.3
47.1 ± 0.1
69.0 ± 0.2
Appendix
Table 4: Instruction tuning Qwen3-8B-Base on Tulu-3 personas-IF on a 32-client ring under a matched budget of 3 epoch. Mean ± std over three seeds; all numbers in %. The first-order and DZO baselines train an FFA-LoRA adapter ( r=8 ); SeedFlood trains the full parameter set through seed-encoded perturbations. Comm. is the total payload one client sends over the 159-iteration run. Learning rates were selected by validation loss from a per-method grid. Best result in bold.
Figure 10: Consensus dynamics of a single gradient under gossip-based model averaging (a) and flooding-based gradient dissemination (b).
1. Overall runtime (ms)
Method
GE
MA
Total time
MeZO
1077
1432
2509
Ours
914
28
942
2. Gradient estimation phase (ms)
Method
Forward pass
Perturbation
Local update
Appendix
Table 5: Detailed wall clock time per iteration of the S eed F lood framework with MeZO versus our subspace coordinate-wise gradient estimator. Results are averaged over 5 steps on OPT-2.7B with batch size 16 and 16 clients (i.e., 16 ZO-gradient messages generated per iteration). “GE” denotes the gradient estimation phase, and “MA” denotes the message-applying phase.
Experiment
Hyperparameter
Value
DSGD
batch size
8
learning rate
{ 1e-4, 1e-5, 1e-6 }
local iteration
5
DSGD (LoRA)
batch size
8
learning rate
{ 1e-2, 1e-3, 1e-4 }
local iteration
5
Appendix
Table 6: Hyperparameter setup for experiment Section 5.2 and Section 5.3 Validation loss is evaluated every one-tenth of the total training iterations, and the model achieving the best validation loss is selected for evaluation on a held-out test set.
Ring Network
OPT-1.3B
Qwen3-1.7B
Cost
Type
Method
SST2
RTE
BoolQ
WiC
MultiRC
ReCoRD
SQuAD
DROP
OPT
Qwen
Δ DSGD
Rel. ZS
-
ZeroShot
53.56
53.43
45.50
56.90
45.40
70.50
57.57
41.11
–
–
-24.29%
100.00%
FO
DSGD
93.69
71.84
66.10
62.33
68.50
71.50
88.59
48.86
526.3GB
681.4GB
0.00%
136.15%
DSGD-LoRA
93.58
70.76
65.70
61.44
68.80
71.20
86.93
49.06
629.1MB
635.8MB
-0.64%
135.30%
DSGD-FFA
92.78
70.04
67.10
60.19
66.50
71.80
86.27
48.31
314.6MB
272.5MB
-1.46%
134.16%
Appendix
Table 7: Generalized performance comparison over a 16-client network among baseline decentralized methods and S eed F lood. SST2–ReCoRD are evaluated on OPT-1.3B and SQuAD/DROP on Qwen3-1.7B. FO and ZO denote first-order and zeroth-order optimization, and the -FFA suffix denotes FFA-LoRA fine-tuning. Cost is the total transmitted volume per edge over training, reported separately for each model. Δ DSGD is the average relative performance difference across all tasks with respect to DSGD (FFT) under the same network topology, and Rel. ZS is the average performance relative to zero-shot. Best results among ZO methods are in bold.
Task / Metric
Method
16
32
64
128
SST2
DSGD-LoRA
86.07
85.35
73.47
55.19
DSGD-FFA
86.58
85.55
85.32
84.44
Choco-LoRA
69.53
55.04
55.16
53.59
Choco-FFA
85.21
84.63
83.68
82.38
S eed F lood
84.57
85.20
84.93
84.82
RTE
DSGD-LoRA
58.36
54.99
51.96
53.67
Appendix
Table 8: Three-run averaged results for the ring-network scalability experiment in Section 5.3 . SST2, RTE, and BoolQ use OPT-125m; SQuAD uses Qwen3-0.6B. Average relative performance is computed with respect to the corresponding reference performance used in the main paper.
Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often distributed across multiple clients, decentralized fine-tuning offers a natural paradigm for collaborative adaptation without a central server. However, enabling full-parameter fine-tuning (FPFT) in this decentralized setting is difficult: FPFT provides strong adaptation capacity but incurs prohibitive resource consumption for billion-scale models. Existing decentralized LLM fine-tuning methods therefore mainly rely on parameter-efficient updates, which improve efficiency but may restrict downstream performance. Moreover, client data are typically non-IID, making decentralized optimization more vulnerable to client drift and unstable convergence. To address these challenges, we propose DECA, a resource-efficient decentralized FPFT framework for LLMs on non-IID data. DECA partitions model parameters into disjoint blocks and performs sequential block-wise Adam optimization, reducing resource consumption while preserving decentralized full-parameter adaptation. To stabilize training, DECA further introduces first- and second-order block-wise moment estimates with fresh local gradient statistics and consensus-derived discrepancy signals. We provide rigorous theoretical analysis and extensive experiments, showing that DECA achieves fast convergence, strong downstream performance, and significant resource efficiency.
Yunsheng Yuan, Shaowei Li, Kai Wang +5
School of Computer Science and Technology, Shandong University, Qingdao China · School of Mathematical Science, Peking University, China · IEIT SYSTEM, China +2
Parameter-efficient fine-tuning methods such as LoRA have become a standard approach for adapting large foundation models. Adopting fine-tuning to distributed settings faces several challenges. Most existing distributed LoRA methods rely on centralized aggregation, and gossip-based decentralized LoRA requires repeated synchronization among multiple model copies. Both methods incur significant communication overhead and introduce errors due to simultaneous aggregation of multiple model updates. In this paper, we take a different perspective and propose a random-walk-based LoRA fine-tuning scheme. Instead of maintaining multiple model replicas, a single model token traverses the network and is updated sequentially using local fine-tuning objectives. This design eliminates the need for global synchronization, substantially reduces communication and computation costs, and avoids aggregation errors. We provide rigorous convergence guarantees for non-convex objectives under standard assumptions. Through empirical results on multiple NLP tasks and graph topologies, we show that the proposed method achieves competitive task performance with substantially less communication and computation than gossip-based LoRA.
Xingran Chen, Rohit Bhagat, Ghadir Ayache +3
Engineering Systems and Design Pillar, Singapore University of Technology and Design, Singapore 487372 · Department of Electrical and Computer Engineering, Rutgers University, Piscataway Township, NJ 08854, USA · LinkedIn, New York, NY 10118, USA +2
Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods require transmitting large collections of trainable weights, assume aligned architectures, and rely on white-box access to model parameters. As model sizes continue to grow and deployments become increasingly heterogeneous, these assumptions become progressively misaligned with practical constraints. We consider an alternative formulation in which collaboration is mediated through model behavior rather than parameters. Clients fine-tune local models on private data and exchange generated outputs on a shared, public prompt set. The server maps these outputs into a semantic representation space, forms a per-prompt semantic consensus, and returns pseudo-labels for further local fine-tuning. This formulation fundamentally changes the communication scaling of federated LLM fine-tuning. The amount of information exchanged depends only on the public prompt budget and the size of the communicated behaviors, independent of model size. As a consequence, the protocol naturally accommodates heterogeneous architectures and applies directly to open-ended text generation. We present a theoretical analysis and empirical results demonstrating that this approach can match strong federated fine-tuning baselines while substantially reducing communication by orders of magnitude (e.g., analytically by a factor of 1006 for Llama3.1-405B), as well as reductions in runtime and energy consumption. These results suggest that, for generative foundation models, behavior-level consensus provides a more appropriate abstraction for federated adaptation than parameter aggregation.
Amr Abourayya, Jens Kleesiek, Michael Kamp
Lamarr Institute for ML and AI, Technical University Dortmund · Institute for AI in medicine (IKIM),University Hospital Essen