Organizations: University of Shanghai for Science and Technology, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China · Zhejiang University, Hangzhou, China
Training large language models (LLMs) at scale incurs substantial communication overhead, while static gradient compression cannot adapt to gradient evolution and may degrade model quality. We propose EDGC, an entropy-driven dynamic gradient compression framework that adapts compression ranks to gradient entropy during training. EDGC combines efficient entropy estimation through gradient sampling, a theoretical model relating entropy to compression rank under a bounded-error constraint, and window-based rank adjustment across pipeline stages. Experiments on 32-V100 and 64-H100 GPU clusters training GPT2 models with 2.5B and 12.1B parameters show that EDGC reduces communication latency by up to 46.45% and end-to-end training time by 16.13%, while maintaining model quality.
Figures & tables
Figure 1: Entropy changes during iteration.
Figure 2: Gradient distribution of GPT2-345M across different model layers at different iterations.
Figure 3: Gradient matrix correlation of GPT2.
Figure 4: Overview of EDGC.
Figure 5: Visualization of warm-up, entropy calculations, rank bounds, and window-based compression adjustments.
Figure 6: Communication time vs. rank values.
Figure 7: Changes in compression error vs. rank values.
Figure 8: The change in loss over time.
Model
Metric
Megatron-LM
PowerSGD
Optimus-CC
EDGC
GPT2-2.5B
Time (day)
18.44
18.14
18.01
15.74
PPL
17.95
22.37
17.97
17.95
GPT2-12.1B
Time (day)
6.88
-
6.42
5.77
PPL
8.73
-
8.84
8.87
Table 1: Training time and PPL after 230K iterations
Tasks
GPT2-2.5B
GPT2-12.1B
Megatron
Pow-SGD
Opti-CC
EDGC
Mega
Opti-CC
EDGC
ARC_easy
41.75%
41.67%
41.73%
40.95%
60.65%
59.84%
58.25%
ARC_challenge
21.84%
21.08%
21.84%
21.76%
28.84%
26.59%
27.47%
HellaSwag
35.55%
30.24%
35.41%
35.64%
42.34%
42.01%
40.65%
OpenBookQA
20.00%
17.40%
20.80%
21.20%
24.60%
25.01%
23.40%
PIQA
65.23%
63.28%
65.23%
65.23%
72.25%
70.50%
70.35%
Table 2: Accuracies on zero-shot tasks
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Cluster 1
Node
Number
8
CPU
Xeon Platinum 8163, 32 cores
Memory
256 GB
GPU
4 × Nvidia Tesla V100 (32 GB)
Interconnect
Intra-node
NVLink (300 Gbps / GPU)
Inter-node
Ethernet (32 Gbps)
Cluster 2
Node
Number
16
Appendix
Table 3: Experimental environments
Figure 9: Gradient entropy under different gradient sampling rates.
Figure 10: Relative change rate of entropy under different iteration sampling rates.
β =1.0
β =0.5
β =0.25
β =0.05
Calculation Time (ms)
78.14
58.7
46.85
39.93
Appendix
Table 4: Time cost under different GSRs ( β )
Figure 11: PPL under different compression strategies.
Method
No Compression
Rank=64
Rank=32
Rank=16
CQM (Ours)
Time (h)
3.0417
3.0167
1.4833
0.7417
1.8750
Appendix
Table 5: Communication time over 30,000 training steps
Metric
Model
w=1
w=100
w=500
w=1000
w=2500
CC
BERT
1.0000
0.9891
0.9838
0.9433
0.8004
GPT-2
1.0000
0.9979
0.9920
0.9807
0.9634
MSE
BERT
0.0000
0.0059
0.0534
0.2742
1.2181
GPT-2
0.0000
0.0066
0.0288
0.1561
0.8750
Appendix
Table 6: Performance metrics (CC, MSE) for varying time window sizes( w )
Figure 12: Effect of stage-aligned rank adaptation on compression error.
Communication has emerged as a critical bottleneck in the distributed training of large language models (LLMs). While numerous approaches have been proposed to reduce communication overhead, the potential of lossless compression has remained largely underexplored since compression and decompression typically consume larger overheads than the benefits of reduced communication traffic. We observe that the communication data, including activations, gradients and parameters, during training often follows a near-Gaussian distribution, which is a key feature for data compression. Thus, we introduce ZipCCL, a lossless compressed communication library of collectives for LLM training. ZipCCL is equipped with our novel techniques: (1) theoretically grounded exponent coding that exploits the Gaussian distribution of LLM tensors to accelerate compression without expensive online statistics, (2) GPU-optimized compression and decompression kernels that carefully design memory access patterns and pipeline using communication-aware data layout, and (3) adaptive communication strategies that dynamically switch collective operations based on workload patterns and system characteristics. Evaluated on a 64-GPU cluster using both mixture-of-experts and dense transformer models, ZipCCL reduces communication time by up to 1.35× and achieves end-to-end training speedups of up to 1.18× without any impact on model quality.
Wenxiang Lin, Xinglin Pan, Ruibo Fan +2
Harbin Institute of Technology, Shenzhen, China · The Hong Kong University of Science and Technology (Guangzhou), China · The Hong Kong University of Science and Technology, Hong Kong SAR
In this paper, we introduce layer-wise curriculum learning for efficient LLM compression. The proposed method facilitates the knowledge transfer from the teacher model to the student model, utilizing a curriculum learning approach that begins with easier optimization tasks and progressively tackles harder ones. In order to adopt the layer-wise learning in LLM compression, we partition the whole model into multiple segments consisting of layers, thereby enabling more computationally efficient knowledge transfer for LLMs. Based on our theoretical analysis of cumulative error phenomenon, layer-wise curriculum learning accelerates convergence while stabilizing the knowledge transfer process. In addition, we present a feature caching method with a multi-threading strategy to efficiently address feature misalignment across layers, maximizing GPU utilization. Consequently, our method exhibits advanced model compression performance, as well as high computational efficiency in terms of minimized memory usage and short training hours. Experiments on multiple datasets show that the proposed method achieves state-of-the-art performance while reducing GPU memory usage and training hours by more than 50% on BERT and GPT-2. Moreover, it outperforms the other pruning methods on LLaMA-family and Qwen models under the same training hours, with a lower GPU memory footprint.
Training large language models (LLMs) is highly memory-intensive, as training must store not only weights and optimizer states but also intermediate activations for backpropagation. While existing memory-efficient methods largely focus on gradients and optimizer states, activation compression is less well established due to the lack of LLM-tailored theory and guarantees. In this work, we develop a theoretical framework showing that activation compression is safe for linear operators when activation compression is unbiased, but problematic for nonlinear ones. We further derive gradient variance bound and establish convergence guarantees for applying activation compression to all linear operators under the standard L-smoothness assumption, showing that it does not change the convergence rate. Guided by the theory, we propose an activation-gradient co-compression method that reuses low-rank activation factors to compress linear-layer gradients without extra computation or additional gradient error. We conduct extensive experiments on Qwen and LLaMA models using a pretraining benchmark and multiple fine-tuning benchmarks to validate our theory and demonstrate competitive performance of our method in both accuracy and compression efficiency. We provide our code in the supplementary material for reproducibility.
Wen-Da Wei, Han-Bin Fang, Yang-Di Liu +3
Nanjing University, Nanjing, China · Tsinghua University, Beijing, China · Huazhong University of Science and Technology, Wuhan, China +1