AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization
Authors: Pengyu He, Yan Zhang, Ruien Li, Guangwen Yang
Organizations: Department of Electrical Engineering, California Institute of Technology · Department of Computer Science and Technology, Tsinghua University · Department of Computer Sciences, University of Wisconsin - Madison
The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state, the communication frequency is reduced while maintaining the training performance. In this work, we introduce AutoLoCo, an adaptive training framework to reduce communication in LLM training. It adapts the local interval using scalar training statistics and corrects each outer update. Our method is motivated by two observations: 1) the appropriate local interval varies across training stages, and 2) changing the number of inner steps per interval creates a mismatch with an unchanged outer optimizer, requiring a correction to the outer update. We optimize this mismatch by correction of the outer optimizer for the momentum and the learning rate using the accumulated inner learning rate. Our experiments under communication constraints demonstrate that AutoLoCo reduces communication frequency by 27% relative to DiLoCo while maintaining training performance.
Figures & tables
Figure 1: AutoLoCo adapts the local horizon and corrects the outer optimizer in both loops.
Method
Training loss ↓
Number of syncs ↓
Payload per sync
Validation loss ↓
DDP
2.7227
50,000†
1×
2.7233
DiLoCo
2.8004
100
1×
2.8020
Linear Interval ‡
2.8706
100
1×
2.8694
QSR ‡
2.8914
41
1×
2.8793
AutoLoCo
2.7843
73
1×
2.7842
Table 1: Quality and communication cost at step 50,000. Training loss averages eight workers over the final 1,000 steps; validation evaluates the final synchronized checkpoint. Payload is measured in full-model BF16 tensors per worker. † counts gradient synchronizations; other counts are outer synchronizations. ‡ Given the differences in training settings, we retain their proxies for selecting communication intervals and adapt the methods to LLM training. Bold and underline indicate the best and second-best results, respectively, in the loss and synchronization count columns.
Loss threshold
3.50
3.20
3.00
2.90
2.85
2.80
Method
Cumulative recorded step time (h) ↓
DiLoCo
1.31
2.83
7.48
10.60
12.44
15.26
Linear Interval
1.68
4.34
9.13
12.87
–
–
QSR
1.69
4.30
9.40
14.27
–
–
AutoLoCo
1.39
3.03
7.52
10.61
12.09
14.22
Table 2: Cumulative recorded step time (hours) to first reach each threshold of the 8 worker mean loss, using a 1,000-step trailing average. A dash indicates that the threshold is not reached within 50,000 steps. The best time in each column is shown in bold.
Method
DDP
DiLoCo
AutoLoCo
Train loss ↓
0.8736
0.9102
0.8759
Syncs ↓
6,490
65
49
Logical payload (GB)
Per sync
–
16.061
16.061
Total ↓
–
1,043.934
786.966
UltraChat test
Table 3: UltraChat training and communication with eight workers, a global batch of 32 for 1 epoch. Training loss covers the final 1,000 steps. Logical payload represents the volume of model data exchanged at the algorithmic level. Actual traffic over physical links depends on the distributed framework and the collective communication algorithm.
Outer correction
Variant
Adaptive intervals
Momentum
Learning rate
Training loss ↓
Syncs ↓
Training PPL
DiLoCo
–
–
–
2.8251
100
16.8618
Adaptive intervals
✓
–
–
2.8310
73
16.9618
Adaptive + momentum
✓
✓
–
2.8484
73
17.2603
AutoLoCo
✓
✓
✓
2.8116
73
16.6362
Table 7: Component ablation on C4 with 4 workers and 50,000 optimizer steps. Adaptive intervals include horizon selection and token mapping. Bold marks the best training loss.
Normalization
Metric
Without
With
Training loss
2.811582
2.826429
Syncs
73
73
Training time (h)
23.1001
23.1300
Table 8: Normalization in C4 pretraining (4 workers)
Normalization
Metric
Without
With
Training loss
2.811582
2.826429
Syncs
73
73
Training time (h)
23.1001
23.1300
Table 8: Normalization in C4 pretraining (4 workers)
Normalization
Metric
Without
With
Training loss
0.876554
0.876079
Syncs
49
49
Test NLL
0.877551
0.877184
Test PPL
2.405003
2.404119
Table 9: Normalization in fine-tuning(16 workers).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Interval assessment
Range of ztmax
Increase supported
ztmax<1.5
Reference consistent
1.5≤ztmax<2.5
Moderate deviation
2.5≤ztmax<4.0
Severe deviation
ztmax≥4.0
Appendix
Table 10: Interval assessments for the active Dt and Ct classifier in the pre-training experiments.
Setting
C4 pre-training
UltraChat SFT
Model
LLaMA (215M total)
Llama-3.1-8B Base
Initialization
Shared random initialization
Shared pretrained checkpoint
Training budget
50,000 steps
One epoch after filtering
Workers
4 / 8
8 / 16
Effective batch per worker
128
4 / 2
Global batch
512 / 1,024
32
Appendix
Table 11: Training settings. Batch sizes count sequences per optimizer step. For SFT, the paired values correspond to eight and sixteen workers. The C4 ablations use four workers and a global batch of 512, with the remaining training settings unchanged.
Setting
C4
UltraChat
Base token horizon Hbase
500
100
Token-horizon bounds [Hmin,Hmax]
[250,750]
[50,150]
Horizon quantum qH
50
10
Base outer LR ηbaseout
0.7
0.7
Base outer momentum μbase
0.7 / 0.9
0.7
Shared setting
Value
Appendix
Table 12: AutoLoCo configuration. Shared settings apply to both tasks.
Block
Cost per round
Preparation + local drift statistics
1,126.07 ms
Aggregated gradient norm for Ct
111.81 ms
Outer preparation
100.70 ms
Post-optimizer scalar/control span
1,090.98 ms
Total
2,429.55 ms
Scalar contributions (worker payload)
256 B
Appendix
Table 13: Control accounting over 73 AutoLoCo rounds. Each block uses the maximum worker duration per round. Scalar bytes are included in endpoint accounting.
Logical payload (GB)
UltraChat test
Method
Train loss ↓
Syncs ↓
Per sync
Total
NLL ↓
PPL ↓
DiLoCo (INT8)
0.9004
65
8.038
522.477
0.9016
2.4636
AutoLoCo (INT8)
0.8761
49
8.038
393.867
0.8772
2.4040
Appendix
Table 14: INT8 pseudo-gradient communication with 16 workers on UltraChat. Optimizer settings follow Appendices E.1 – E.3 . INT8 applies to pseudo-gradient communication.
Communication-efficient pre-training of LLMs is increasingly important as training draws on compute distributed across clusters, data centers, and lower-bandwidth links. Many practical methods reduce communication frequency but still rely on synchronous All-Reduce operations that maintain identical model states and tie progress to global collectives. This can become a bottleneck when bandwidth or worker speed is heterogeneous. We introduce GASLoC, a novel decentralized pre-training algorithm that generalizes the notion of communication acceleration to the recently popular "outer optimizer" to allow a practical gossip-based training framework that is compatible with adaptive optimizers, allows for local optimizer steps, and can utilize sparse randomized peer communication. Empirically, on a number of standard LLM training tasks, we demonstrate that GASLoC outperforms state-of-the-art decentralized algorithms in single step per communication setting for a number of topologies and, unlike existing decentralized methods in the LLM setting, it allows to obtain performance competitive with DiLoCo when utilizing multiple local steps. In the heterogeneous bandwidth setting we demonstrate the advantage of GASLoC showing that it can significantly outperform DiLoCo.
Pietro Cagnasso, Eugene Belilovsky, Edouard Oyallon
Concordia University · Mila · CNRS, Sorbonne University
Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect. Scaling up these clusters is expensive and can become impractical, imposing limits on the size of models that can be trained. Several recent studies have proposed training methods that are less communication intensive, avoiding the need for compute clusters with extremely high interconnect speeds. These low communication training methods still employ a global synchronization step for model parameters, which can be too costly with a high number of participants, as the communication cost scales quadratically with group size. In this work, we propose a novel optimization method, NoLoCo, that does not explicitly synchronize all model parameters during training and does not require any collective communication. NoLoCo implicitly synchronizes model weights via a novel variant of the Nesterov momentum optimizer by partially averaging model weights within randomly selected subgroups. We provide both a theoretical convergence analysis of our optimizer and empirical results from language model training. Our method requires significantly less communication than fully sharded data parallel training and DiLoCo, a widely used low-communication baseline. Moreover, our method avoids global blocking communication, thereby reducing accelerator idle time. Our experiments show that NoLoCo is more communication-efficient than DiLoCo, improving final perplexity by up to 4% and converging up to 4× faster in wall-clock time across a range of worker counts, model sizes, and communication bandwidths.
Low-Rank Adaptation (LoRA) has gained popularity as a fine-tuning approach for Large Language Models (LLMs) due to its low resource requirements and good performance. While numerous studies have investigated ways to improve LoRA serving efficiency by serving multiple LoRAs concurrently, existing methods assume that a wide range of LoRA adapters are available for serving. In our work, we conduct extensive empirical studies to show that current LoRA training paradigms do not efficiently utilize hardware resources and incur high overhead to obtain a performant LoRA adapter. Leveraging these insights, we propose PLoRA, which automatically orchestrates concurrent LoRA fine-tuning jobs under given hardware and model constraints and develops performant kernels to improve training efficiency. Across a range of LLMs and LoRA configurations, PLoRA improves training throughput by up to 12.8x and reduces the overall fine-tuning makespan by up to 7.52x compared to existing approaches.
Minghao Yan, Zhuang Wang, Zhen Jia +2
University of Wisconsin-Madison · Amazon Web Services