Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.
Figures & tables
Fig. 1: The core problem of weight tying that breaks ghost clipping’s assumption.
Model
Setting: SGD
Accuracy (%)
GPU Memory (MB)
Trainable Params
DistilGPT2
WT
86.77±0.76
3676.64
81,912,576
No-WT
86.77±0.76
3824.64
120,509,952
GPT2
WT
89.09±0.69
5390.99
124,439,808
No-WT
89.09±0.69
5538.99
163,037,184
TABLE I: Effect of Weight Tying under Standard Non-Private Training on SST-2. Weight tying preserves predictive utility while improving parameter efficiency, reducing GPU memory usage, and slightly shortening runtime across both decoder-only models.
Model
Setting: DP-SGD with normal clipping
Accuracy (%)
Peak GPU Memory (MB)
DistilGPT2
WT
80.61±1.32
13813.55
No-WT
83.23±0.25
14256.41
GPT2
WT
78.44±1.73
18530.09
No-WT
81.48±1.07
18973.67
TABLE II: Effect of Weight Tying under DP-SGD with normal clipping on SST-2. Untying embeddings consistently improves utility and optimization stability under DP across both DistilGPT2 and GPT2.
Model
Clipping Strategy (No WT)
Accuracy (%)
Peak GPU Memory (MB)
DistilGPT2
ghost clipping
83.23±0.25
4747.19
normal clipping
83.23±0.25
14256.41
GPT2
ghost clipping
81.48±1.07
6786.90
normal clipping
81.48±1.07
18973.67
TABLE III: Comparison between ghost clipping and standard DP-SGD normal clipping under untied embeddings (No-WT) on SST-2. When the additive gradient separability assumption remains valid, ghost clipping preserves identical utility and optimization stability while substantially reducing GPU memory consumption across both DistilGPT2 and GPT2.
TABLE IV: Comparison of weight tying and clipping strategies under DP-SGD on SST-2 (DistilGPT2). The best utility is achieved with untied weights, while the most efficient configuration in terms of memory is Ghost Clipping with untied weights.
TABLE V: Comparison of weight tying and clipping strategies under DP-SGD on SST-2 (GPT2). The best utility is achieved with untied weights, while the most efficient configuration in terms of memory is Ghost Clipping with untied weights.
Training Setting
Embedding
DistilGPT2
GPT2
Accuracy (%)
Std. Dev.
Accuracy (%)
Std. Dev.
Standard SGD
WT
86.77
0.76
89.09
0.69
No-WT
86.77
0.76
89.09
0.69
DP-SGD (Normal)
WT
80.61
1.32
78.44
1.73
No-WT
83.23
0.25
81.48
1.07
DP-SGD (Ghost)
No-WT
83.23
0.25
81.48
1.07
TABLE VI: Overall comparison of standard and differentially private training under weight tying (WT) and untied embeddings (No-WT) on SST-2 For DistilGPT2 and GPT2. Best DP utility is shown in bold.
Dataset
Task Type
Split
Size (MB)
#Samples
SST-2
Sentiment Analysis
Train
11.82
67,349 (59,560 bal.)
Validation
0.15
872
Test
0.32
1,821
QQP
Paraphrase Detection
Train
63.85
363,846 (268,756 bal.)
Validation
7.09
40,430
Test
68.61
390,965
TABLE VII: GLUE Dataset Statistics with Task Types and Balanced Training Sizes
TABLE VIII: Comparison of weight tying and clipping strategies under DP-SGD on QNLI (DistilGPT2). Untied embeddings improve utility, while Ghost Clipping preserves these gains with substantially lower memory usage.
TABLE IX: Comparison of weight tying and clipping strategies under DP-SGD on QNLI (GPT2). Untied embeddings improve utility, while Ghost Clipping preserves these gains with substantially lower memory usage.
TABLE X: Comparison of weight tying and clipping strategies under DP-SGD on QQP (DistilGPT2). Untied embeddings improve utility, while Ghost Clipping preserves these gains with substantially lower memory usage.
TABLE XI: Comparison of weight tying and clipping strategies under DP-SGD on QQP (GPT2). Untied embeddings improve utility, while Ghost Clipping preserves these gains with substantially lower memory usage..
TABLE XII: Comparison of weight tying and clipping strategies under DP-SGD on SST-2 (DistilGPT2) with ϵ=1 . The best utility is achieved with untied weights, while the most efficient configuration in terms of memory is Ghost Clipping with untied weights.
TABLE XIII: Comparison of weight tying and clipping strategies under DP-SGD on SST-2 (DistilGPT2) with ϵ=5 . The best utility is achieved with untied weights, while the most efficient configuration in terms of memory is Ghost Clipping with untied weights.
Large language models (LLMs) are trained on vast datasets that may contain sensitive information. Differential privacy (DP), the de facto standard for formal privacy guarantees, provides a principled framework for training LLMs with provable privacy protection. However, state-of-the-art DP training implementations rely on fast gradient clipping techniques with memory overhead O(Bmin{T2,d2}), where B is the batch size, T is the sequence length, and d is the model width. This becomes prohibitive as both model size and context length grow. We propose DP-SGD-RC, a novel variant of DP-SGD with randomized clipping that reduces memory and compute complexity. DP-SGD-RC leverages stochastic trace estimation methods, specifically Hutchinson's estimator[Hutchinson, 1989] and its improved variant, Hutch++[Meyer et al., 2021], to reduce the memory footprint of per-sample gradient norm estimation. We provide a tight privacy analysis showing that DP-SGD-RC achieves noise multipliers competitive with deterministic clipping. Experiments fine-tuning Llama~3.2-1B on long-context benchmarks spanning classification, question answering, and summarization tasks demonstrate that DP-SGD-RC matches baseline utility while significantly reducing memory and compute requirements.
Enayat Ullah, Sai Aparna Aketi, Devansh Gupta +2
Meta Platforms Inc. · University of Southern California
Large language models (LLMs) are commonly adapted to downstream tasks through fine-tuning, but fine-tuning data often contains sensitive information that may be leaked by the resulting model. Differential privacy (DP) offers formal protection against such leakage, yet DP fine-tuning of LLMs still suffers from substantial utility degradation due to gradient clipping and noise injection. Existing work improves this trade-off by combining DP with parameter-efficient fine-tuning methods such as LoRA, which constrain the form of updates. In this work, we study a complementary direction: selective fine-tuning, which constrains where updates are applied. We propose DP-SelFT, a framework for differentially private selective fine-tuning of LLMs. DP-SelFT addresses three DP-specific challenges in parameter selection: avoiding repeated privacy cost, improving stability under noisy estimates, and selecting parameters that remain useful under clipped and noisy updates. It first constructs a lightweight DP synthetic dataset and performs selection only on this synthetic data, so the selection stage incurs no additional privacy cost. It then conducts layer-level selection by temporarily training candidate layer subsets on a synthetic training split and evaluating them on a synthetic validation split. Crucially, this temporary training is performed under a perturbation regime matched to downstream DP fine-tuning, with worst-case perturbations of the same scale as DP noise. This favors layer subsets that are not only learnable but also robust to noisy private updates. Experiments on benchmark tasks show that DP-SelFT consistently improves the privacy--utility trade-off over existing DP fine-tuning baselines under the same privacy guarantees.
Haichao Sha, Zihao Wang, Yuncheng Wu +2
Renmin University of China · Nanyang Technological University
Federated learning (FL) enables the collaborative training of large-scale language models (LLMs) across edge devices while keeping user data on-device. However, FL still exposes sensitive information through client-provided gradients. Differentially private stochastic gradient descent (DP-SGD) mitigates this risk by clipping each client's contribution to a threshold C and adding noise proportional to C. Existing adaptive clipping techniques dynamically adjust C but demand tedious hyperparameter tuning, which can erode the privacy budget. In this paper, we introduce DP-LAC, a method that first estimates an initial clipping threshold within an order of magnitude of the optimum using private histogram estimation, and then adapts this threshold during training without consuming additional privacy budget or introducing new hyperparameters. Empirical results show that DP-LAC outperforms both state-of-the-art adaptive clipping methods and vanilla DP-SGD, achieving an average accuracy gain of 6.6%.
Haaris Mehmood, Jie Xu, Karthikeyan Saravanan +2
Samsung R&D Institute UK (SRUK) · Samsung AI Centre Cambridge