Authors: Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir
Organizations: El Oued University, El Oued, Algeria · King Fahd University of Petroleum and Minerals (KFUPM), Dhahran, Saudi Arabia · Interdisciplinary Research Center For Smart Mobility and Logistics, KFUPM, Dhahran, Saudi Arabia
Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to a unique recursive error phenomenon known as the Bellman drift. This article introduces the Homomorphic Advantage Operator (HAO), a stabilization framework designed to prevent polynomial approximation divergence in FHE-based deep RL. HAO adapts the zero-mean centering projection from advantage-based value estimation directly to temporal-difference (TD) targets. This linear projection annihilates the uniform state-value baseline that drives the Bellman drift, maintaining per-state action rankings while requiring zero additional non-linear multiplicative depth and avoiding expensive ciphertext bootstrapping. The proposed HAO framework was evaluated using a three-tier experimental methodology, including a tabular Markov Decision Process (MDP), an encrypted CartPole environment using real CKKS cryptographic operations, and a 20-node logistics routing benchmark with dense continuous features. The results demonstrate that the proposed HAO strictly bounds network pre-activations within the safe polynomial approximation domain. The proposed HAO RL agents achieved 0% boundary breaches across all random seeds used, whereas regularization alone (L2 weight decay and gradient clipping) breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes. Finally, HAO agents improve optimal policy accuracy by 18.0 percentage points in tabular domains and remain stable when DP-SGD-style Gaussian noise is added to the clipped gradients.
Figures & tables
#
Solution
Reference
HE
Paradigm
Mechanism Summary
Training-Time Stabilization (Supervised)
1
Boundary Loss
Araki et al. [ 3 ]
LHE
Supervised
Exponential penalty on ∣z∣>B ; prevents explosion during training
Table 1: Taxonomy of existing solutions for polynomial activation and privacy-preserving learning under HE. PHE = Partially HE, LHE = Leveled HE (no bootstrapping), FHE = Fully HE (with bootstrapping).
Figure 1: High-level overview of the HAO framework for privacy-preserving reinforcement learning. The system integrates three interdependent components: (1) CKKS-encrypted state transmission and polynomial forward evaluation, (2) the HAO centering projection that neutralizes Bellman drift, and (3) clipped weight updates with optional DP-SGD-style Gaussian noise.
Figure 2: HAO client-server architecture for privacy-preserving RL. The client encrypts private states via CKKS, the server computes an encrypted forward pass with polynomial activations, and the client applies HAO centering after decryption. Weight updates can be perturbed with DP-SGD-style Gaussian noise. The HAO centering (gold) is the core contribution: it removes the V(s) Bellman drift without any ciphertext–ciphertext multiplication.
Experiment
Environment
Architecture
Training Parameters
Evaluated Modes
1: Tabular MDP
10 states, 4 actions
1 Hidden Layer (244 params)
2000 epochs, 5 random seeds
Standard, Full HAO, Centering Only
2: Encrypted CartPole
Hybrid client-server using real TenSEAL CKKS operations
Full HAO, Centering Only, Reg Only, No Stab., HAO+DP
Table 2: Summary of Experimental Setup Methodology.
Figure 3: Pre-activation growth over 2000 epochs. Centering successfully bounds polynomial inflation and stabilizes training.
Figure 4: Gradient norm dynamics on a logarithmic scale over 2000 epochs. The uncentered baseline exhibits exponential gradient explosion upon breaching the polynomial boundary, whereas HAO centering compresses gradient variance.
Mode
Accuracy (%)
Max ∥z∥∞
Grad Norm
Standard
68.0± 8.4
1.83±0.20
5.04±0.34
Full HAO
86.0 ± 5.5
0.74 ± 0.07
0.06 ± 0.00
Centering Only
86.0 ± 5.5
0.70 ± 0.08
0.01 ± 0.00
Table 3: Tabular MDP (2000 epochs, 5 seeds)
Figure 5: Trapping region (pre-activation trajectory) for the hybrid TenSEAL CKKS CartPole agent. The network utilizes a single hidden layer architecture.
Figure 6: Optimization health (gradient norm) over 250 episodes (log scale). Standard LHE diverges exponentially.
Mode
Max ∥z∥∞
Max ∥∇θ∥
Max.
Avg.
Reward
Q ⋅ P time (s)
Standard LHE
28.68
25,003
28.0
N/A
HAO LHE
3.61
2.88
93.0
1.64
Table 4: CKKS Encrypted CartPole (250 episodes). Q ⋅ P time is the mean time per episode of the encrypted centering step.
Figure 7: Pre-activation growth over 1,000 episodes. The HAO framework strictly bounds pre-activations below B=5.0 , whereas uncentered baseline agents diverge under the environment shift.
Figure 8: Agent performance (moving average reward) over 1,000 episodes in the 20-node Logistics environment. Full HAO maintains optimal performance post-shift.
We present the first theoretical convergence analysis of machine learning training under fully homomorphic encryption (FHE), combined with a differentially private (DP) training algorithm tailored to encrypted computation. Our approach improves computational efficiency over standard differentially private gradient descent (DP-GD) while achieving comparable utility. In particular, we prove convergence of approximate gradient descent using polynomial approximations of activation and loss functions, which are required for FHE compatibility. To preserve privacy in downstream tasks, we integrate differential privacy without relying on costly per-sample gradient clipping, enabling scalable encrypted learning. We also provide data-independent hyperparameter selection and theoretically grounded strategies for polynomial approximation which can be of independent interest. Together, these contributions advance the feasibility of efficient, private, and secure machine learning on sensitive data.
Yvonne Zhou, Mingyu Liang, Ivan Brugere +5
University of Maryland, College Park, MD 20742 · J.P. Morgan AI Research, New York, NY, 10017 · AlgoCRYPT CoE
Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplications may be composed only to a bounded depth before a costly bootstrapping operation is required to continue. Every nonlinearity must therefore be approximated by an iterative method; each iteration increasing the number of multiplications. A higher iteration count buys precision but exhausts the available depth more frequently and thus triggers more bootstraps, which dominate latency. We introduce Homomorphic Encryption-Aware Training (HEAT), a fine-tuning method that makes the per-nonlinearity iteration counts learnable, enabling them and the model weights to co-adapt during training. HEAT optimizes iterations with respect to the task objective, allowing the model to adapt to approximation errors encountered during inference without architectural changes or retraining from scratch. We further relate iteration count to quantization bit width and bound, at fixed weights, the gap between our objective and quantization-aware training. On encrypted GPT-2 decoding, HEAT reduces iterations by 3.1×, bootstraps by 1.6×, and end-to-end latency by 1.4×, while improving decode agreement over the calibrated encrypted baseline.
Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos +2
Sapienza University of Rome · George Mason University · Paradigma
Fully homomorphic encryption (FHE) supports only additions and multiplications, so FHE-only neural-network inference typically replaces ReLU with polynomials fitted over empirical activation intervals. Such interval fitting often requires higher-degree polynomials to control activation error, incurring homomorphic evaluation costs, while classification is determined by the final logit decision. We revisit ReLU replacement from a decision-aware perspective: given a trained single-hidden-layer ReLU MLP and a specified calibration set, can an HE-friendly low-degree polynomial replace ReLU without retraining while preserving calibration-set decisions? We focus on quadratic replacement, the lowest-degree that retains a genuine per-unit nonlinearity. For calibration sets positive-margin separable in the lifted space, we formulate quadratic replacement as a linear separation problem, yielding necessary and sufficient conditions for calibration-lossless replacement and a constructive algorithm for the coefficients. When the positive-margin condition fails -- often because a few near-boundary or misclassified calibration samples bring the lifted hulls into contact -- we extend the same geometric framework via reduced convex hulls and Lagrangian-dual soft-margin relaxations. These cap the weight any single sample can carry, converting the problem into smaller convex quadratic programs that yield approximately feasible coefficients with high empirical agreement on calibration-set decisions. In particular, at the maximal weight cap μ=1, the reduced-convex-hull relaxation reduces to standard convex-hull separation; the relaxation thus continuously extends the positive-margin exact theory. Under CKKS, the quadratic replacement matches plaintext top-1 accuracy on multiple benchmarks, running 3.7--4.1× faster than Remez-7 in the activation module and 1.18--1.68× faster end-to-end.
Rui Li, Wenyuan Wu, Weijie Miao
Chongqing Key Laboratory of Secure Computing for Biology, Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences, 266 Fangzheng Avenue, Beibei District, Chongqing 400714, China · Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University, Hung Hom, Hong Kong, China