Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift
Authors: Marcel Mordarski, Benjamin Gras, Abdelrahman Shehata, Daniel Budina, Roberto Bondesan
Organizations: Department of Computing, Imperial College London, London SW7 2AZ, United Kingdom · Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ, United Kingdom
Quantum key distribution (QKD) links are provisioned from security analyses of stationary channels, whereas the devices that determine the channel drift between recalibrations. Whether an eavesdropper who cannot alter the channel's own noise gains by following that drift has not been quantified. Adaptive eavesdropping is posed here as a constrained Markov decision process in which the attacker selects one circuit per round while the noise level follows an Ornstein--Uhlenbeck process and the abort condition is a budget over each block of rounds. The value of adaptation is bounded by the best fixed circuit and a dynamic-programming upper bound. The actions are learnt attacks. Whereas Decker et al. trained a parametrised circuit on a fixed gate template against a fixed channel, here the gate structure and rotation angles are searched jointly. This yields circuits compact enough to form a discrete action set, extending the construction to noise models lacking a known template, including the amplitude damping channel. On device-independent E91 under bilateral depolarising noise, a reinforcement-learning attacker raises her Holevo information from 0.135 for the best fixed circuit to 0.348 at zero detection, 98% of the upper bound. On BB84 under a drifting bit-flip channel, she exceeds a conservative noise-indexed rule by 0.024 in fidelity, reaching 99% of the upper bound. Under stationary noise, the attacker's gain from basis asymmetry changes sign between an averaged and a per-basis error-rate constraint. The search, started from random gate sequences, recovers the analytical cloners and the collective-attack key rate, and meets the lower bound of the Winick--Lütkenhaus--Coles objective from above.
Figures & tables
Figure 1: Attainable attacks under stationary channel noise at an abort threshold of FAB⩾0.87 . (a) The maximum attainable FAE under the trusted-noise convention as the channel’s total honest error weight, ε , increases. This is plotted for two scenarios: a channel concentrating all its error in a single basis, and one splitting the error evenly. Because both channels produce the same observed disturbance, the noiseless closed-form solution (dashed line) assigns them identical values. The shaded region highlights the gap between this theoretical limit and the actual attainable FAE . Each data point represents the best outcome from three random seeds, with a cross-seed spread of approximately 0.001 at ε=0.10 (Sec. S4 ). (b) The attacker’s gain from basis asymmetry under two different monitoring conventions. When constrained by the basis-averaged quantum bit error rate (QBER), the gain is small and positive. However, when constraints are applied to the per-basis rates separately (shown here under the untrusted convention where the attacker must reproduce each observed rate), the gain changes sign and increases by more than an order of magnitude. The trusted per-basis values in Table 1 exhibit this same sign change.
ε
constraint
Noise
Symmetric FAE
One-basis FAE
Gain
0.10
averaged
trusted
0.7846
0.7858
+0.0012
0.10
per-basis
trusted ∗
0.7846
0.7631
−0.0214
0.10
per-basis
untrusted
0.7220
0.6737
−0.0483
0.22
averaged
trusted
0.6581
0.6640
+0.0059
0.22
per-basis
trusted ∗
0.6581
abort
—
0.22
per-basis
untrusted
0.8153
0.7301
−0.0853
Table 1: Basis asymmetry evaluated across two feasibility constraints and two noise conventions. Under the averaged constraint, the attacker can redistribute the disturbance between the bases, provided that FAB⩾0.87 . Under the trusted per-basis constraint, each observed error rate must independently remain at or below 0.13 . These specific rows (marked ∗ ) are evaluated using Eq. ( 4 ) – a result that the search independently reproduces in the averaged rows. Notably, at ε=0.22 , the one-basis channel already fails this per-basis test. Under the untrusted per-basis constraint, the attacker supplies the entire disturbance and must reproduce each observed rate to within a tolerance of 2×10−3 . Finally, absolute values should only be compared between rows that share the same noise convention.
Component
BB84
E91/DIQKD
Protocol family
Prepare-and-measure
Entanglement-based, CHSH game
Channel
Bit-flip of rate pbf , before the attacker’s interaction
Bilateral depolarising ΛpA⊗ΛpB , after the preparation
Attacker’s memory
One qubit
One qubit (three-qubit preparation on A , B , E )
Honest reference
FAB=1−pbf/2
Sobs=(1−p)2⋅22
Abort threshold (episode mean)
FAB⩾0.87
Sobs⩾τS≈2.444
Attacker’s information
FAE , memory measured in the announced basis
χ(A:E)
Table 2: Component-by-component correspondence between the two drift environments. The monitored statistic is averaged over each episode of T=50 rounds, and the drift is the discretised Ornstein–Uhlenbeck process of Eq. ( 8 ) with one round as the time step. The cliff is the noise level at which the undisturbed channel alone reaches the abort threshold for BB84 and the zero-key-rate value Ssec for E91 (Eq. ( 7 )).
Figure 2: Disturbance versus information in the device-independent setting. At p=0 , the discovered attacks trace the noiseless Acín envelope [ 1 , 41 ] , demonstrating that the search successfully saturates the analytical frontier where one exists. However, when averaged over the drifting noise, the attainable value (marked by the star) drops to χ≈0.35 , falling short of the 0.607 permitted by the theoretical envelope at the threshold. This gap occurs for two reasons: the theoretical envelope assumes a strictly noiseless channel ( p=0 ) while the attainable value averages over active noise ( p>0 ), and the constructed attacks are restricted to a fixed three-qubit preparation.
Policy
FAE/χ
Det.
% of VDP
BB84
Conservative lookup
0.7197
0%
96
Greedy oracle
0.7434
0%
99
Adaptive, strict
0.7434
0%
99
Adaptive, tolerated
0.7477
4%
99
Dyn. prog., VDP
0.7527
0%
100
E91
Best fixed circuit
0.135
0%
38
Table 3: Value of adaptation. Mean information per round and detection rate, alongside percentages of the zero-detection dynamic-programming upper bound, VDP , evaluated over the same library. BB84 values are averaged over 50 evaluation episodes. E91 agent and oracle values are averaged over 200 episodes, while its fixed-circuit and dynamic-programming baselines are evaluated over 2000 episodes. The E91 upper bound is resolved computationally to a grid precision of 0.001 (Sec. S8 ). The BB84 conservative lookup policy is not a fixed circuit; it dynamically selects the safest library level anchored nearest the current noise level.
Figure 3: Adaptive attack on BB84 under a drifting device. Upper panel: attacker information FAE during evaluation, reaching the greedy oracle (dashed) at its best checkpoints and well above the passive baseline (dotted). Lower panel: detection rate, which stays at zero at those checkpoints.
Figure 4: BB84 under a drifting bit-flip channel: final evaluation FAE across policies, strict-feasibility variants only. The value-based agent matches the greedy oracle at 0.7434 , unmasked policy-gradient methods plateau around 0.02 below it, and the noise-indexed conservative lookup (“Static safe”) reaches 0.7197 .
Comparison
Mean difference
95% CI
Wilcoxon p
Adaptive − greedy oracle
−0.0023
[−0.0028,−0.0018]
2.5×10−15
Adaptive − masked policy gradient
+0.0017
[+0.0012,+0.0022]
1.0×10−8
Adaptive − dynamic programming (7% det.)
−0.0056
[−0.0065,−0.0047]
1.3×10−21
Table 4: Paired per-episode comparisons on E91, n=200 episodes with shared channel seeds. Two-sided Wilcoxon signed-rank p -values [ 57 ] .
Figure 5: E91 under the bilateral depolarising channel: training curve for the adaptive attack. Upper panel, Holevo information during evaluation against the greedy oracle (dashed), the coarse-grid dynamic-programming run (dotted) and the passive baseline (grey); lower panel, detection rate.
Figure 6: E91/DIQKD: final evaluation χ across policies. The masked value-based agent reaches 0.3484 against a greedy oracle at 0.3507 and a coarse-grid dynamic-programming run at 0.3540 with 7% detection, whereas unmasked policy-gradient methods collapse to χ≈0.26 – 0.28 by confining themselves to three or four distinct actions.
Inner optimiser
FAE
FAB
Gates
Evaluations
Best gates
CMA-ES
0.764±0.001
0.794±0.000
5.2±0.7
2000
4
Stochastic perturbation (SPSA) [ 51 ]
0.760±0.003
0.794±0.002
6.4±0.8
2000
6
Uniform random
0.756±0.006
0.796±0.002
5.6±1.0
4000
6
Table S1: Inner-loop optimiser comparison on the BB84 cloning task; mean ± standard deviation over five seeds. All three recover FAE to within 0.01 of one another and reach the same equilibrium of four to six gates, so the discovered circuit size is set by the budget and is insensitive to the choice of algorithm.
Threshold τ
FAB attained
FAE attained
Closed form
Ratio
0.95
0.9500
0.7179
0.7179
1.0000
0.90
0.9000
0.8000
0.8000
1.0000
0.87
0.8700
0.8363
0.8363
1.0000
0.85
0.8500
0.8571
0.8571
1.0000
Table S2: Validation on the noiseless symmetric channel. The search receives no analytical input and recovers the closed form FAE=21+FAB(1−FAB) at every threshold tested.
Q
χ attained
h(Q)
r⩽1−h(Q)−χ
1−2h(Q)
0.02
0.1414
0.1414
0.7171
0.7171
0.04
0.2423
0.2423
0.5155
0.5154
0.06
0.3274
0.3274
0.3451
0.3451
0.11
0.4999
0.4999
0.0002
0.0002
Table S3: Key-rate upper bounds from constructed attacks on the symmetric channel. The search maximises χ behind the same feasibility constraint and is given no analytical input; it returns χ=h(Q) at every Q tested, so the upper bound it produces coincides with the symmetric Shor–Preskill rate 1−2h(Q) [ 49 ] to within 10−4 .
ε
Honest FAB
Symmetric
One-basis
Asymmetry gain
Closed-form excess
0.02
0.990
0.8278
0.8278
<0.0001
+0.0085
0.06
0.970
0.8081
0.8087
+0.0006
+0.0276
0.10
0.950
0.7846
0.7858
+0.0012
+0.0505
0.16
0.920
0.7366
0.7400
+0.0034
+0.0963
0.22
0.890
0.6581
0.6640
+0.0059
+0.1723
Table S4: Attainable FAE against total honest channel error ε at threshold FAB⩾0.87 , for a channel placing all error in one basis ( px=ε ) and one splitting it evenly ( px=pz=ε/2 ). Both present honest fidelity 1−ε/2 , so the noiseless closed form evaluated at the observed FAB assigns both the same value, 0.8363 . The final column is the excess of that closed form over the attainable value under the trusted-noise convention, and measures the conservatism of the closed form for a physically restricted attacker.
γ
Honest FAB
FAE
Gates
Ancilla
FAE
Gates
0.05
0.9812
0.8197
5
1 qubit
0.7850
6
0.10
0.9622
0.7997
4
2 qubits
0.7858
8
0.15
0.9430
0.7751
4
3 qubits
0.7858
5
0.20
0.9236
0.7437
5
Table S5: Amplitude damping, a non-Pauli channel outside the Pauli-channel cloning solutions of Refs. [ 10 , 30 ] , and attacker ancilla size at the one-basis operating point ( ε=0.10 ). Both at threshold FAB⩾0.87 .
Level
Dimension
Unknowns
S⩽
Solve time
1
5
13
2.828427
0.02 s
1+AB
9
25
2.828427
0.03 s
2
13
41
2.828427
0.05 s
Table S6: The relaxation reproduces Tsirelson’s bound 22 [ 13 ] at three monomial sets, which validates the implementation. The dimension is the side of the moment matrix, and the unknowns are the number of distinct moments. Solved with SCS [ 40 ] at ε=10−9 on one core.
S
Pguess⩽
Analytical
Residual
Solve time
2.828
0.51229
0.51229
6.9×10−10
0.25 s
2.750
0.66536
0.66536
1.9×10−12
0.05 s
2.600
0.77839
0.77839
7.3×10−12
0.08 s
2.444
0.85592
0.85592
8.8×10−11
0.04 s
2.200
0.94441
0.94441
5.0×10−11
0.05 s
Table S7: Device-independent guessing probability at level 1+AB , with two moment-matrix blocks of dimension nine. Agreement with the closed form holds at the solver tolerance, which validates the implementation.
S
NPA via min-entropy
Acín envelope
Attained by construction
2.828
0.0350
0.0021
—
2.600
0.6386
0.4184
—
2.444
0.7755
0.6069
≈0.350
Table S8: Upper and lower bounds on the eavesdropper’s Holevo information. The first two columns hold without reference to Hilbert-space dimension, the first through the min-entropy and the second as the tight collective-attack bound on the von Neumann quantity, and the third is exhibited by a circuit. The attained value is averaged over the operating noise distribution, which is why it sits furthest below.
level 1+AB
level 2
m
dim
moments
time
dim
moments
time
2
9
25
0.10 s
13
41
0.06 s
3
16
100
0.14 s
28
244
0.24 s
4
25
289
0.33 s
49
865
0.80 s
5
36
676
0.59 s
76
2276
2.17 s
6
49
1369
1.20 s
109
4969
8.66 s
Table S9: Growth of the relaxation with the number of measurement settings per party, at two monomial sets. Dimension is the side of the moment matrix. The circuit search costs 43 – 49 s per operating point independently of m , since it optimises over circuits and not over correlations.
QZ
QX
χ(AZ:E)
h(QX)
Gates
r⩽
r⩾
Gap
0.0501
0.0519
0.2941
0.2944
11
0.4191
0.4188
3×10−4
0.0684
0.0320
0.2043
0.2043
5
0.4356
0.4356
<10−4
0.0895
0.0120
0.0938
0.0938
4
0.4716
0.4716
<10−4
0.0992
0.0020
0.0202
0.0204
6
0.5135
0.5132
3×10−4
Table S10: Constructed attacks against the reliable numerical lower bound of Ref. [ 58 ] at matched per-basis error rates. Attacks use a two-qubit ancilla and are matched to the per-basis rates shown. χ(AZ:E) is the information the circuit holds about the key-basis bit and h(QX) the largest value any attack could hold at that phase error rate. The final columns bound the asymptotic key rate from above, as 1−h(QZ)−χ(AZ:E) for the circuit found, and from below by the numerical minimisation.
DTU Electro, Technical University of Denmark, DK-2800, Kgs. Lyngby, Denmark · Nokia Bell Labs, 91300 Massy, France · Centre for Quantum Optical Technologies, CeNT, University of Warsaw, 02-097 Warszawa, Poland +1