Pareto-optimal quantum kernel selection for unsupervised anomaly detection on real malware beaconing data
Authors: Boaz Micah, Nadia Milazzo, Maissa Beji, Borja Aizpurua, Llorenç Espinosa-Portalés, Esteban Payares, Ghada Ben Slama, Luc Andrea, +3 more
Organizations: Multiverse Computing, Parque Científico y Tecnológico de Gipuzkoa, Paseo de Miramón 170, Planta 2, 20014 Donostia / San Sebastián, Spain · IQM Quantum Computers, 4 rue Royale, 75008 Paris, France · Multiverse Computing, 7 rue de la Croix Martre, 91120 Palaiseau, Paris, France · Department of Basic Sciences, Tecnun – University of Navarra, San Sebastián, Spain · IQM Quantum Computers, Georg-Brauchle-Ring 23-25, 80992 Munich, Germany · Allianz Quantum Hub, Paris, France
Quantum kernel methods are leading candidates for a practical quantum advantage in machine learning, but assessing that potential requires two quantities usually reported separately: how well a kernel performs on the task, and how far its geometry departs from the classical kernels available for the same problem. We introduce a fully unsupervised, multi-objective protocol that optimises simultaneously the normalised pseudo discrepancy (NPD), a label-free proxy for anomaly detection quality, and the geometric difference (GD) to a tuned classical reference kernel, selecting models from the resulting Pareto front. We apply it to malware beaconing detection in real network traffic, using a one-class support vector machine with fidelity and projected quantum kernels over four data encodings, on simulators and on IQM's 20-qubit Garnet processor. NPD-guided selection alone finds a fidelity kernel that beats the tuned classical baseline, but with a geometric difference too small to certify the gain as quantum. Projected kernels reach far larger geometric differences; the Pareto-selected one only marginally exceeds the baseline (AUC 0.782 versus 0.765, gC→Q≈89>N relative to that reference kernel), still below the NPD-selected fidelity kernel (0.840).
Figures & tables
Model
Hyperparameter
Search Space
Projected Quantum Kernel (RBF)
Feature map
Z, IQP, CNOT, Hamiltonian
Number of layers
{1,2,3,4,5}
K-RDM
{1}
Gamma ( γ )
[100μ,10]
ν
[0,1]
Fidelity Kernel
Feature map
Z, IQP, CNOT, Hamiltonian
Table 1: Hyperparameter search space for quantum kernel methods. The description of the considered feature maps is given in Appendix A . The number of layers ranged from 1 to 5. We only considered the single-qubit reduced density matrix for the projected quantum kernels, and γ is the hyperparameter of the projected quantum kernel. ν is a hyperparameter of the OCSVM model, which represents an upper bound on the fraction of training errors and a lower bound of the fraction of support vectors. Curly brackets {} denote discrete values, while square brackets [] denote continuous ranges.
Figure 1: The unsupervised optimisation pipeline: a hyperparameter search over quantum kernels selected under three objectives (NPD, GD, and their joint Pareto knee) and evaluated on the test set.
Feature Name
Forward packets per second
Backward packets per second
Standard deviation of the FFT of forward traffic
Standard deviation of forward inter-arrival time
Standard deviation of packet lengths
Mean forward packet length
Table 2: Feature set obtained from the feature engineering procedure
AUC
F1 (A)
F1 (N)
0.765
0.599
0.561
Table 3: Performance of the classical OCSVM on the malware dataset.
NPD
GD
Pareto
Feature map
Hamiltonian
Hamiltonian
Hamiltonian
Layers
2
1
5
ν
0.322
–
0.266
gC→Q
4.297
7.14
5.94
NPD
2.34
0.592
0.72
Table 4: Best fidelity quantum kernel configuration found for the three objectives on the dataset using the statevector simulator. The GD objective depends only on the kernel matrix, so ν is not defined for that column (–).
NPD
GD
Pareto
Feature map
Z
Hamiltonian
IQP-style
Layers
5
1
5
ν
0.385
–
0.104
γ
2.37
9.99
0.000395
gC→Q
32.0
272
88.6
NPD
2.94
0.603
1.42
Table 5: Best projected quantum kernel configurations found for the three objectives on the six features dataset using the statevector simulator. The GD objective depends only on the kernel matrix, so ν is not defined for that column (–).
Figure 2: Pareto fronts of the fidelity (blue) and projected (red) quantum kernels under joint maximisation of the NPD and GD, evaluated on the statevector simulator. Circles mark the non-dominated configurations of each front; stars mark the selected knee points. The fidelity kernel maintains a low GD across the front, whereas the projected kernel trades a high GD at low NPD for improved NPD, exposing the tension between the two objectives. Each front contains only the configurations that remained non-dominated during the optimisation, which is why the two kernels contribute different numbers of points.
Figure 3: Test AUC versus geometric difference gk for all hyperparameter configurations evaluated during the GD-maximisation run (statevector simulation). Blue: fidelity kernels; red: projected quantum kernels. The green dashed line marks gk=N≈14.14 , below which a classical kernel is guaranteed to match the quantum model; the orange dotted line is the tuned classical OCSVM baseline (AUC =0.765 ). Fidelity kernels stay below the threshold yet reach the highest AUC in the run ( ≈0.84 ), whereas projected kernels extend to gk∼270 with an AUC that shows no systematic trend with gk . A few projected configurations combine gk≫N with an AUC above the baseline, but none reaches the best fidelity kernels, and the two largest-GD models (yellow, green) perform at or below the baseline.
Figure 4: Test AUC versus NPD for all hyperparameter configurations evaluated during the NPD-maximisation run (statevector simulation). Blue: fidelity kernels; red: projected quantum kernels. The orange dotted line marks the tuned classical OCSVM baseline (AUC =0.765 ). Fidelity kernels concentrate at higher AUC, with a substantial fraction above the baseline, while projected kernels spread widely between AUC ≈0.33 and 0.86 and mostly lie below it. Within each family the NPD–AUC relation is weak. The largest-NPD configuration of each family (yellow: fidelity; green: projected) lies above the baseline, at AUC =0.840 and 0.777 respectively.
Method
NPD
GD
Pareto (NPD & GD)
AUC
F1(A)
F1(N)
AUC
F1(A)
F1(N)
AUC
F1(A)
F1(N)
Statevector Simulator
Fidelity
0.840
0.657
0.715
0.666
0.494
0.800
0.812
0.671
0.773
Projected
0.777
0.599
0.563
0.738
0.542
0.340
0.782
0.638
0.736
IQM Noisy Simulator
Fidelity
0.790
0.600
0.582
0.752
0.582
0.630
0.798
0.600
0.576
Table 6: Performance of the optimised quantum kernels obtained from the optimisation of the three objectives: NPD, GD and Pareto frontier using the statevector simulator. We re-run the optimal hyperparameters selected by each optimisation using the IQM noisy simulator; the OCSVM parameter ν was additionally re-tuned on the noisy simulator by maximising the NPD. The classical OCSVM baseline achieves AUC =0.765 . The best value of each metric is shown in bold.
Figure 5: Layout of the IQM Garnet QPU. The color map indicates single and two-qubit gate errors (darker=higher error, lighter=lower error). We highlight the optimal way to choose a qubit patch to implement a Hamiltonian evolution feature map with nearest-neighbour interactions. The specific patch needs to be adjusted depending on re-calibration of the QPU.
Method
Feature Map
AUC
F1(A)
F1(N)
Fidelity
Hamiltonian
0.532
0.640
0.151
Projected
Z
0.612
0.630
0.656
Table 7: Performance of the quantum kernels obtained from the IQM Garnet QPU using the optimal configurations for the hyperparameters as given by the NPD score; for the projected kernel, ν was selected to maximise the test scores.
Global College, Shanghai Jiao Tong University, Shanghai 200240, China · School of Electrical and Computer Engineering, Cornell University, Ithaca, New York 14850, United States · Shanghai Jiao Tong University, Shanghai 200240, China +1