Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storing full gradients is prohibitively expensive at LLM scale. We study how to compress these gradients while preserving influence estimates for future queries that are unknown at storage time. Through a worst-case analysis, we characterize the optimal fixed-dimensional linear representation and propose eigenbasis-corrected one-bit gradient projection (EOGP) to approximate it at scale. Specifically, EOGP uses EK-FAC to reduce gradient dimensionality, then applies PCA within the retained subspace to learn compression directions from the training gradients. We then apply one-bit quantization to the resulting coordinates, allowing more coordinates to be retained within a fixed storage budget. On GPT-2, EOGP predicts retraining outcomes more accurately than the evaluated compression baselines while using one-sixteenth of their per-example storage. On OLMo 2 SFT models from 1B to 32B parameters, EOGP remains competitive with the baselines allocated over 100 times as much storage per example.
Figures & tables
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Source checkpoint
Dataset
GPT-2
gpt2
WikiText-2
OLMo 2 1B
allenai/OLMo-2-0425-1B-SFT
Tülu 3 mixture
OLMo 2 7B
allenai/OLMo-2-1124-7B-SFT
Tülu 3 mixture
OLMo 2 13B
allenai/OLMo-2-1124-13B-SFT
Tülu 3 mixture
OLMo 2 32B
allenai/OLMo-2-0325-32B-SFT
Tülu 3 mixture
Appendix
Table 1: Source model checkpoints and datasets.
Model
Uncompressed FP16 (GB/example)
One-bit representation (KB/example)
OLMo 2 1B
2.147
28.896
OLMo 2 7B
12.952
57.792
OLMo 2 13B
25.376
72.240
OLMo 2 32B
62.411
115.584
Appendix
Table 2: Per-example representation payloads at ku=2,048 per module. Uncompressed gradients cover the attributed parameters, and one-bit representations include packed signs and FP16 scales. Shared artifacts are accounted for in Section E.1 .
Model
EOGP PCA fitting (min)
Store construction (s/example)
EOGP
LoGra
OLMo 2 1B
4.5
0.0563
0.0577
OLMo 2 7B
8.9
0.0999
0.1315
OLMo 2 13B
11.1
0.1474
0.1661
OLMo 2 32B
18.4
0.2470
0.2780
Appendix
Table 3: PCA fitting and store-construction times on one NVIDIA B200. Fitting starts from cached coordinates and includes saving the correction matrices. Store-construction totals combine separately timed stages, including disk I/O.
Method
12 KB
24 KB
96 KB
EOGP
0.902
0.890
0.919
LoGra (Random)
0.604
0.447
0.086
LoGra (PCA)
0.750
0.570
−0.355
GraSS
0.330
0.079
−0.205
LoRIF
0.820
0.823
0.822
Appendix
Table 4: Pearson correlation between FP16 and one-bit influence scores on GPT-2. We report the median of the query-wise correlations over 481 queries. Storage budgets refer to one-bit storage per training example. The highest correlation at each budget is shown in bold.
Training large language models (LLMs) at scale incurs substantial communication overhead, while static gradient compression cannot adapt to gradient evolution and may degrade model quality. We propose EDGC, an entropy-driven dynamic gradient compression framework that adapts compression ranks to gradient entropy during training. EDGC combines efficient entropy estimation through gradient sampling, a theoretical model relating entropy to compression rank under a bounded-error constraint, and window-based rank adjustment across pipeline stages. Experiments on 32-V100 and 64-H100 GPU clusters training GPT2 models with 2.5B and 12.1B parameters show that EDGC reduces communication latency by up to 46.45% and end-to-end training time by 16.13%, while maintaining model quality.
Qingao Yi, Jiaang Duan, Jun Zhang +5
University of Shanghai for Science and Technology, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China · Zhejiang University, Hangzhou, China
Training large language models (LLMs) is highly memory-intensive, as training must store not only weights and optimizer states but also intermediate activations for backpropagation. While existing memory-efficient methods largely focus on gradients and optimizer states, activation compression is less well established due to the lack of LLM-tailored theory and guarantees. In this work, we develop a theoretical framework showing that activation compression is safe for linear operators when activation compression is unbiased, but problematic for nonlinear ones. We further derive gradient variance bound and establish convergence guarantees for applying activation compression to all linear operators under the standard L-smoothness assumption, showing that it does not change the convergence rate. Guided by the theory, we propose an activation-gradient co-compression method that reuses low-rank activation factors to compress linear-layer gradients without extra computation or additional gradient error. We conduct extensive experiments on Qwen and LLaMA models using a pretraining benchmark and multiple fine-tuning benchmarks to validate our theory and demonstrate competitive performance of our method in both accuracy and compression efficiency. We provide our code in the supplementary material for reproducibility.
Wen-Da Wei, Han-Bin Fang, Yang-Di Liu +3
Nanjing University, Nanjing, China · Tsinghua University, Beijing, China · Huazhong University of Science and Technology, Wuhan, China +1
Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective for 4-bit activations and 8-bit gradients, which would easily cause slow convergence or accuracy loss. To address this, we introduce AGoQ, incorporating two new techniques: 1) a layer-aware activation quantization algorithm that allocates appropriate bit-widths for activations of various layers based on their types and pipeline stages to achieve near 4-bit activation storage, and 2) a gradient quantization algorithm that reduces memory usage and shortens communication time by employing 8-bit gradient storage and precision-preserving 8-bit All-Reduce communication. We conduct extensive experiments using different sizes of LLMs on two GPU clusters (up to 64 GPUs), and the experimental results show that our AGoQ reduces the memory by up to 52% and achieves up to 1.34× improvement of training speed compared to state-of-the-art training systems Megatron-LM (w/ or w/o ZeRO), COAT and DeepSpeed with 8B to 32B LLaMA models, while achieving convergence loss on pretraining and comparable accuracy on downstream tasks with LLaMA architectures.
Wenxiang Lin, Juntao Huang, Luhan Zhang +5
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen · Huawei Technologies Ltd.