HE-OFT: Privacy-Preserving One-Shot Federated Fine-Tuning under Homomorphic Encryption
Authors: Halil İbrahim Kanpak, Sinem Sav, Alptekin Küpçü
Organizations: Department of Computer Engineering, Kocı University, İstanbul, Türkiye · Department of Computer Engineering, Bilkent University, Ankara, Türkiye
Many organizations adapt large pretrained models to their own tasks by fine-tuning on private data. Several of these parties often hold data for the same task and wish to fine-tune a model together without pooling that data. Federated learning (FL) enables joint fine-tuning, but reconstruction attacks on shared intermediate values (the model or its gradients) remain a privacy risk. A one-shot protocol that exchanges one encrypted contribution exposes no intermediate value. Such a protocol still gives the trained model to every participant, which is not permitted where the model is a regulated or proprietary asset. We present HE-OFT, the first cryptographically secure one-shot federated fine-tuning protocol in which no party receives the trained model. Each client fine-tunes a low-rank adapter and a classifier head on a frozen public backbone and keeps the adapter. The client uploads one encrypted head displacement, which the server combines under multiparty CKKS and never decrypts. A quorum of clients returns only the predicted label to the querier. On four text classification tasks and one vision task, HE-OFT reaches 61 to 79 per cent accuracy, against 20 to 48 per cent for a client training alone. HE-OFT keeps 85 to 96 per cent of the accuracy of a disclosed model. A test-time query takes 443.1 to 1713.1 s on one core, or 56.1 to 255.1 s with level restoration on a GPU. Restoring levels at the server cuts the traffic per query from up to 1.6 GiB to 13.5 MiB.
Figures & tables
Figure 1: HE-OFT in three stages: local fine-tuning, aggregation and selection, and test-time query serving. Each client fine-tunes an adapter and a classifier head on the same frozen public backbone. The client keeps the adapter and uploads one encrypted head displacement weighted by its per-class counts. The server adds the ciphertexts and divides by the per-class totals under encryption. The shared head θ⋆ exists only as a ciphertext.
Symbol
Meaning
N , C , d
numbers of clients and classes, feature dimension
φ , φj
public backbone, and backbone with client j ’s adapter Aj
θ0 , hj , Δj
public head initializer, client j ’s head, Δj=hj−θ0
gj
client j ’s per-class counts, gj,c=nj,c
θ⋆
shared head, held only as a ciphertext
Enc
encryption under the collective public key
Table 1: Notations.
Figure 2: Choosing between two arrangements without decrypting either. Each client scores both on data held back from its own training. One product with the client’s encrypted one-hot label selects the true-class logit. The per-class counts of correct predictions are weighted by each class’s share of the federation’s examples, taken from the count matrix already used for the aggregation. The shared arrangement pools evidence across clients because it is one model, and the personal arrangement does not. Only the classes a client holds add to its score, which is thus a lower bound. A quorum decrypts one value only.
Figure 3: Serving one test-time query. The querier computes the features with its own backbone and adapter and encrypts them. The server thus never sees the feature vector and cannot invert it to recover the query. The server holds the evaluation keys only. It applies the encrypted head, obtains encrypted logits, and reduces them to the index of the largest entry by a tournament of homomorphic comparisons. A quorum of clients then re-encrypts that index under the querier’s key. Quorum members other than the querier hence contribute shares without learning the index.
Figure 4: HE-OFT along four axes. Panels (a) to (c) vary the federation size, the label skew and the local trajectory on DBpedia around N=10 , α=0.1 and K=200 . Panel (d) gives the test time of one query at N=10 against the number of classes, returning the label or the scores. Times are single-core measurements or GPU projections ( Section V-E ). Panels (a) and (c) are one seed and panel (b) is the mean over three.
servable
reference
Task
C
shared
personal
alone
disclosed
pooled
head
adapter
AG-News
4
0.649
0.479
0.475
0.739
0.921
TREC
6
0.607
0.429
0.400
0.711
0.967
DBpedia
14
0.789
0.754
0.451
0.925
0.988
Banking77
77
0.206
0.686
0.249
0.760
0.923
Table 2: Accuracy on the global test set, three seeds, N=10 , α=0.1 . The servable columns are the two arrangements from which HE-OFT answers queries, and bold marks the better one. The reference columns cannot be served under the threat model of Section IV . Alone is one client’s own model without federation, disclosed is the merged model that HE-OFT does not build, and pooled is centralized training on the union of the data. On one seed the Dirichlet draw leaves some clients with fewer than twenty samples, too few to train. The effective federation is thus seven on AG-News and nine on TREC.
Task
shared head
personal adapter
alone
disclosed
AG-News
0.649 ( 0.217 )
0.479 ( 0.021 )
0.475 ( 0.006 )
0.739 ( 0.073 )
TREC
0.607 ( 0.074 )
0.429 ( 0.068 )
0.400 ( 0.051 )
0.711 ( 0.043 )
DBpedia
0.789 ( 0.019 )
0.754 ( 0.014 )
0.450 ( 0.009 )
0.925 ( 0.020 )
Banking77
0.206 ( 0.030 )
0.686 ( 0.012 )
0.249 ( 0.005 )
0.760 ( 0.018 )
CIFAR-100
0.748 ( 0.010 )
0.774 ( 0.003 )
0.197 ( 0.009 )
0.784 ( 0.004 )
Table 3: Standard deviation over the three seeds of Table 2 , in parentheses beside each mean. The spread is below 0.03 everywhere except AG-News and TREC. On AG-News the shared head reaches 0.809 , 0.736 and 0.402 , and the third partition drops three clients.
Label skew α
0.05
0.10
0.30
1.00
selected arrangement
0.814
0.789
0.915
0.970
selected
shared
shared
personal
personal
Clients N
10
20
50
selected arrangement
0.803
0.826
0.882
Table 4: Sensitivity on DBpedia along the protocol axes, each varied around the default of N=10 , α=0.1 and 200 local steps. The row reports the accuracy of the arrangement the estimator selects. The skew row is the mean over three seeds. The local-steps and client-count rows are a single seed, and their default cells are that seed’s value rather than the three-seed mean.
Figure 5: Accuracy of the arrangement each selection rule serves, per task of Table 2 , averaged over three seeds. The black mark is the better of the two servable arrangements, chosen per seed.
Selection rule
correct cells
regret
majority of held-out accuracies
8/27
0.074
same, class-balanced within a client
11/27
0.085
always the shared head
19/27
0.198
global-prior estimator, zero fill
23/27
0.021
global-prior estimator, rare-class fill
24/27
0.020
Table 5: Choosing between the two servable arrangements without decrypting either. Cells are counted over the five tasks and the four CIFAR-10 partitions, at three seeds each. Regret is the mean accuracy a wrong choice gives up, over the cells in which the rule chooses wrongly.
N
α
Aloc
Asel
Adis
Asel/Adis
5
0.10
0.384
0.949
0.963
0.986
5
0.30
0.570
0.952
0.966
0.985
20
0.04
0.212
0.950
0.947
1.003
20
0.16
0.397
0.953
0.961
0.992
Table 6: CIFAR-10 on two published partitions, three seeds, on the whole 10,000 -image test set. The disclosed model aggregates and decrypts both the adapter and the head. The last column is the share of the disclosed model’s accuracy that HE-OFT keeps without disclosing a model. All are measured on HE-OFT at each prior paper’s partition.
Figure 6: Per-operation cost against the two axes that set it, from one measurement over the whole grid. Left: the per-query arithmetic at N=10 , which depends on the ring degree and not on the federation size. Right: the key switch that returns one label, which grows with both. Read in seconds, the left panel spans 0.0007 to 0.18 s and the right panel spans 0.10 to 1.7 s. Neither is the cost of a query. Measured end to end, one query totals 443.1 s at four classes and 1713.1 s at a hundred, of which the encrypted argmax between them is 96 to 99 per cent.
System
Model
Train time
Test time
Output
POSEIDON ( N=10 ) [ 22 ]
3-layer MLP
5,283 s
0.38 s
scores
POSEIDON ( N=50 ) [ 22 ]
CNN
175 h ∗
3.6 s ∗
scores
slytHErin
NN20
–
245.6 s †
scores
CryptPEFT
ViT-B/16
–
2.26 s
scores
HE-OFT ( N=10 ), scores only
RoBERTa-base
4.21 s
CPU 2.6 s GPU 0.7 s
scores
HE-OFT ( N=10 )
RoBERTa-base
4.21 s
CPU 443.1 s GPU 56.1 s
label
Table 7: Cost against the closest systems, as published. Train time is the cost of producing the served model under encryption, and a dash means no training. Test time is the latency of one query, at HE-OFT’s best case of four classes, with GPU figures projected ( Section V-E ). Output is what the querier receives. From scores, d+1=769 answers recover HE-OFT’s head, and from labels a copy of fidelity 0.90 needs 16 to 260 times as many ( Section V-F ). Hardware differs across rows. ∗ Extrapolated. † Per batch.
When
Item
Per client
setup, once
collective public key share
4.25 MiB
relinearization key share
102.0 MiB
rotation keys, seven of them
238.0 MiB
training, once
encrypted head displacements
12 to 44 MiB
per query
encrypted features, uploaded
7.5 MiB
encrypted label, returned
1.0 MiB
Table 8: Communication at N=10 , on the fifteen-modulus chain at ring degree 215 that the encrypted argmax requires. Share sizes do not depend on the number of clients. The per-query total does, because the quorum returns one key-switching share each. It does not depend on the number of classes, because the level restorations are local to the server.
C
4
6
14
77
100
collective refresh, s
31.2
47.3
63.7
112.4
113.0
server bootstrap, s
400.8
612.5
832.4
1473.0
1468.3
bootstraps
17
26
35
62
62
error, server bootstrap
1.3 to 7.3×10−5 , against exact for the refresh
Table 9: The two level-restoration mechanisms at N=10 , evaluation at ring degree 215 . The specified design restores levels at the server under collectively generated bootstrapping keys at ring degree 216 . The measured alternative restores them by a collective refresh, which costs a share from every client. Latency is of the encrypted argmax alone, single-run wall clock on one core of a VALAR CPU node.
Surface
Queries
Fidelity
Membership
model in plaintext [ 18 , 20 , 21 ]
–
1
0.50 to 0.69
scores per query [ 22 , 24 , 25 ]
769
1.000
0.49 to 0.53
updates per round [ 13 ]
–
–
TPR 25.3% at FPR 0.1%
labels only, HE-OFT
1.2×104 to 2.0×105
0.90
0.47 to 0.52
Table 10: What each surface gives an adversary, on HE-OFT’s heads over three tasks, three seeds and both arrangements. The model-in-plaintext row adds CIFAR-100. The per-round row gives its source’s setting and true-positive rate (TPR) at a false-positive rate (FPR) of 0.1 per cent. Queries is the cost of a copy at the fidelity shown, and a dash marks a column that does not apply. Fidelity measures how much of the hidden model an adversary obtains. Membership, the area under the curve of our strongest attack, measures what the surface reveals about the training data, and 0.5 is a random guess.
Figure 7: Fidelity of a copy of the served head against the queries spent per parameter of the head, mean over three seeds and both arrangements. Each line is one task under the label-only interface that HE-OFT serves. Each hollow marker, in the colour and shape of its task, is an interface that returns scores, where solving from d+1 answers recovers the head exactly.
Task
∞
8
4
2
1
0.5
AG-News
fidelity
0.989
0.986
0.954
0.889
0.774
0.598
accuracy
0.649
0.648
0.620
0.494
0.372
0.303
DBpedia
fidelity
0.956
0.946
0.831
0.571
0.268
0.139
accuracy
0.789
0.785
0.640
0.297
0.150
0.103
Banking77
fidelity
0.777
0.722
0.364
0.052
0.020
0.016
accuracy
0.206
0.203
0.092
0.028
0.018
0.015
Table 11: Randomized response on the returned label. Copy fidelity after 105 label queries, against the accuracy the served model retains, at each privacy budget ε in the column heads, where ε=∞ means no randomization. Mean over three seeds. The majority share a constant predictor reaches is 0.488 on AG-News, 0.250 on DBpedia and 0.083 on Banking77.
System
Model
Rounds
Protection
Model in plaintext
Querier receives
DENSE [ 18 ]
ResNet-18
one
none
yes
model
GH-OFL [ 20 ]
VGG-16
one
none
yes
model
FedAUXfdp [ 21 ]
MobileNetV2
one
DP
yes
model
SHE-LoRA [ 23 ]
Llama-3.1-70B
many
partial HE
yes
model
POSEIDON [ 22 ]
CNN
many
HE
no
scores
slytHErin [ 24 ]
NN50
–
HE
no
scores
Table 12: The closest systems on four axes, as each paper reports them. The model column names the largest model each paper evaluates. The plaintext column says whether any party ends up holding the trained model in plaintext, and provider means one party does while the querier does not. A dash in the rounds column marks slytHErin and CryptPEFT, which serve a model they do not train. HE, DP and 2PC are homomorphic encryption, differential privacy and two-party computation. POSEIDON keeps the model under the collective key and states that the querier does not obtain the model’s weights.
Homomorphic encryption (HE) enables privacy-preserving aggregation in federated learning (FL) by allowing the server to operate on encrypted data without decryption. Existing HE-over-the-air (OTA) methods mainly rely on single-key HE schemes and require channel estimation or pre-equalization to compensate for wireless fading. However, single-key HE remains vulnerable to honest-but-curious (HBC) clients holding the shared secret key, while multi-key HE provides stronger client-level security by assigning each device its own secret key. We propose a four-phase protocol that enables the aggregation of xMK-CKKS over a shared wireless channel without channel estimation. The protocol retransmits partial public keys and ciphertexts through the same channel realization, so that the dominant large-modulus encryption terms cancel algebraically during decryption. We integrate this protocol with zero-order FL over slowly varying LoS-dominant channels, where each device transmits a single encrypted scalar per round and the communication/encryption overhead is independent of the model dimension. We show that the residual noise induced by encryption and wireless aggregation preserves the standard convergence rate O(1/K) up to a negligible noise floor, where K is the number of communication rounds. The protocol assumes a non-trusted server and is secure against HBC clients, preventing any client from recovering the local updates of other participants. Numerical results on MNIST and CIFAR-10 validate the theoretical analysis.
Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor. We take a different lever. Mapping networks generate a network's weights from a small trainable latent through a frozen affine projection; because the map is shared and affine, averaging latents is exactly averaging the generated weights. We turn this into a practical low-bandwidth federated channel with two changes: a low-rank, seed-regenerable factorisation of the projection (cutting generator memory from ~80 GB to ~10 MB), and a delta formulation θ=θpre+UV⊤z that learns an additive correction around a shared centrally-pretrained base -- federated fine-tuning, which is what makes the method work at scale. A frozen orthogonal classifier head further removes the head from the payload while improving accuracy. On CIFAR-100 with ResNet-18+GroupNorm, our method (FLITE, Federated Low-rank Iterative Training Engine) communicates 1,280 floats (~5 KB) per client per round -- an 8718x reduction -- and reaches 74.67%, within ~0.5 pp of full-weight FedAvg. The averaging identity holds to floating-point precision (6×10−8); the method sits one to two orders of magnitude below PowerSGD and top-k on the bandwidth-accuracy Pareto; it matches or exceeds full-weight FedAvg under strong non-IID skew. int4 latents reach 648 bytes per round at unchanged accuracy, whereas int4 full-weight FedAvg collapses to chance.
Fine-tuning large language models (LLMs) on domain-specific data is essential for downstream adaptation. In many deployments, a participant cannot hold the complete model locally. This happens because the model owner keeps the full model proprietary, or because the participant lacks sufficient compute resources. Split Learning (SL) addresses this by partitioning the model between the participant and a server so that only a small portion runs locally. When the underlying data is additionally distributed across multiple institutions with privacy requirements, Federated Learning (FL) further enables collaborative training across participants by sharing only model updates instead of raw data. In this combined setting, each client transmits intermediate activations to the server, and for LLM fine-tuning, this exchange poses an inherent privacy paradox. The autoregressive nature of LLMs causes the transmitted activations to leak the input, and existing perturbation-based defenses are fundamentally ineffective in this setting. We address this leakage through a learned obfuscate-and-recover scheme that protects participants' private datasets while still allowing an independently deployable model to be trained on the server side. Experiments demonstrate that our approach achieves strong privacy protection with modest utility loss and system overhead, making split-based federated LLM fine-tuning practically viable.