Organizations: University of Shanghai for Science and Technology, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China · Zhejiang University, Hangzhou, China
Training large language models (LLMs) at scale incurs substantial communication overhead, while static gradient compression cannot adapt to gradient evolution and may degrade model quality. We propose EDGC, an entropy-driven dynamic gradient compression framework that adapts compression ranks to gradient entropy during training. EDGC combines efficient entropy estimation through gradient sampling, a theoretical model relating entropy to compression rank under a bounded-error constraint, and window-based rank adjustment across pipeline stages. Experiments on 32-V100 and 64-H100 GPU clusters training GPT2 models with 2.5B and 12.1B parameters show that EDGC reduces communication latency by up to 46.45% and end-to-end training time by 16.13%, while maintaining model quality.
Figures & tables
Figure 1: Entropy changes during iteration.
Figure 2: Gradient distribution of GPT2-345M across different model layers at different iterations.
Figure 3: Gradient matrix correlation of GPT2.
Figure 4: Overview of EDGC.
Figure 5: Visualization of warm-up, entropy calculations, rank bounds, and window-based compression adjustments.
Figure 6: Communication time vs. rank values.
Figure 7: Changes in compression error vs. rank values.
Figure 8: The change in loss over time.
Model
Metric
Megatron-LM
PowerSGD
Optimus-CC
EDGC
GPT2-2.5B
Time (day)
18.44
18.14
18.01
15.74
PPL
17.95
22.37
17.97
17.95
GPT2-12.1B
Time (day)
6.88
-
6.42
5.77
PPL
8.73
-
8.84
8.87
Table 1: Training time and PPL after 230K iterations
Tasks
GPT2-2.5B
GPT2-12.1B
Megatron
Pow-SGD
Opti-CC
EDGC
Mega
Opti-CC
EDGC
ARC_easy
41.75%
41.67%
41.73%
40.95%
60.65%
59.84%
58.25%
ARC_challenge
21.84%
21.08%
21.84%
21.76%
28.84%
26.59%
27.47%
HellaSwag
35.55%
30.24%
35.41%
35.64%
42.34%
42.01%
40.65%
OpenBookQA
20.00%
17.40%
20.80%
21.20%
24.60%
25.01%
23.40%
PIQA
65.23%
63.28%
65.23%
65.23%
72.25%
70.50%
70.35%
Table 2: Accuracies on zero-shot tasks
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Cluster 1
Node
Number
8
CPU
Xeon Platinum 8163, 32 cores
Memory
256 GB
GPU
4 × Nvidia Tesla V100 (32 GB)
Interconnect
Intra-node
NVLink (300 Gbps / GPU)
Inter-node
Ethernet (32 Gbps)
Cluster 2
Node
Number
16
Appendix
Table 3: Experimental environments
Figure 9: Gradient entropy under different gradient sampling rates.
Figure 10: Relative change rate of entropy under different iteration sampling rates.
β =1.0
β =0.5
β =0.25
β =0.05
Calculation Time (ms)
78.14
58.7
46.85
39.93
Appendix
Table 4: Time cost under different GSRs ( β )
Figure 11: PPL under different compression strategies.
Method
No Compression
Rank=64
Rank=32
Rank=16
CQM (Ours)
Time (h)
3.0417
3.0167
1.4833
0.7417
1.8750
Appendix
Table 5: Communication time over 30,000 training steps
Metric
Model
w=1
w=100
w=500
w=1000
w=2500
CC
BERT
1.0000
0.9891
0.9838
0.9433
0.8004
GPT-2
1.0000
0.9979
0.9920
0.9807
0.9634
MSE
BERT
0.0000
0.0059
0.0534
0.2742
1.2181
GPT-2
0.0000
0.0066
0.0288
0.1561
0.8750
Appendix
Table 6: Performance metrics (CC, MSE) for varying time window sizes( w )
Figure 12: Effect of stage-aligned rank adaptation on compression error.
Harbin Institute of Technology, Shenzhen, China · The Hong Kong University of Science and Technology (Guangzhou), China · The Hong Kong University of Science and Technology, Hong Kong SAR