Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
Organizations: El Oued University, El Oued, Algeria · King Fahd University of Petroleum and Minerals (KFUPM), Dhahran, Saudi Arabia · Interdisciplinary Research Center For Smart Mobility and Logistics, KFUPM, Dhahran, Saudi Arabia
Abstract
Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to a unique recursive error phenomenon known as the Bellman drift. This article introduces the Homomorphic Advantage Operator (HAO), a stabilization framework designed to prevent polynomial approximation divergence in FHE-based deep RL. HAO adapts the zero-mean centering projection from advantage-based value estimation directly to temporal-difference (TD) targets. This linear projection annihilates the uniform state-value baseline that drives the Bellman drift, maintaining per-state action rankings while requiring zero additional non-linear multiplicative depth and avoiding expensive ciphertext bootstrapping. The proposed HAO framework was evaluated using a three-tier experimental methodology, including a tabular Markov Decision Process (MDP), an encrypted CartPole environment using real CKKS cryptographic operations, and a 20-node logistics routing benchmark with dense continuous features. The results demonstrate that the proposed HAO strictly bounds network pre-activations within the safe polynomial approximation domain. The proposed HAO RL agents achieved 0% boundary breaches across all random seeds used, whereas regularization alone (L2 weight decay and gradient clipping) breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes. Finally, HAO agents improve optimal policy accuracy by 18.0 percentage points in tabular domains and remain stable when DP-SGD-style Gaussian noise is added to the clipped gradients.
Figures & tables
| # | Solution | Reference | HE | Paradigm | Mechanism Summary |
| Training-Time Stabilization (Supervised) | |||||
| 1 | Boundary Loss | Araki et al. [ 3 ] | LHE | Supervised | Exponential penalty on ; prevents explosion during training |
| 2 | Selective Grad. Clipping | Araki et al. [ 3 ] | LHE | Supervised | Clips gradients excluding BatchNorm , ; preserves normalization |
| 3 | SLAF | Pulido-Gaytan & Tchernykh [ 27 ] | LHE | Supervised | Trainable polynomial coefficients; learns task-specific activations |
| 4 | HE-Friendly Training | Baruch et al. [ 4 ] | LHE | Supervised | Knowledge distillation + polynomial BN folding |
| 5 | Weight Clamping / L2 | Standard technique | Any | Supervised | Constrains weight norms; bounds pre-activations indirectly |
| Experiment | Environment | Architecture | Training Parameters | Evaluated Modes |
| 1: Tabular MDP | 10 states, 4 actions | 1 Hidden Layer (244 params) | 2000 epochs, 5 random seeds | Standard, Full HAO, Centering Only |
| 2: Encrypted CartPole | Hybrid client-server using real TenSEAL CKKS operations | 1 Hidden Layer (114 params) | 250 episodes | Standard vs. HAO LHE |
| 3: Logistics Ablation | 20-node routing, 8D continuous states, 6 actions, | 1 Hidden Layer (486 params) | 1000 eps, 5 seeds, shift at ep 500 | Full HAO, Centering Only, Reg Only, No Stab., HAO+DP |
| Mode | Accuracy (%) | Max | Grad Norm |
| Standard | 8.4 | ||
| Full HAO | 86.0 5.5 | 0.74 0.07 | 0.06 0.00 |
| Centering Only | 86.0 5.5 | 0.70 0.08 | 0.01 0.00 |
| Mode | Max | Max | Max. | Avg. |
| Reward | Q P time (s) | |||
| Standard LHE | N/A | |||
| HAO LHE | 3.61 | 2.88 | 93.0 | 1.64 |
| Mode | Max | Breach % | FHE-Deploy? |
| Full HAO | 3.89 | 0.0% | Yes |
| Centering Only | 7.79 | 54.5% | No |
| Regularization Only | 6.98 | 23.8% | No |
| No Stabilization | 7.82 | 83.8% | No |
| HAO + DP Moderate | 4.12 | 0.0% | Yes |
| HAO + DP Strong | 4.01 | 0.0% | Yes |
| Reg. Only | 2.14 | 4.91 | Breach | Breach | Breach | Breach |
| Full HAO | 1.25 | 1.87 | 2.95 | 4.88 | Breach | Breach |