Pareto-optimal quantum kernel selection for unsupervised anomaly detection on real malware beaconing data
Authors: Boaz Micah, Nadia Milazzo, Maissa Beji, Borja Aizpurua, Llorenç Espinosa-Portalés, Esteban Payares, Ghada Ben Slama, Luc Andrea, +3 more
Organizations: Multiverse Computing, Parque Científico y Tecnológico de Gipuzkoa, Paseo de Miramón 170, Planta 2, 20014 Donostia / San Sebastián, Spain · IQM Quantum Computers, 4 rue Royale, 75008 Paris, France · Multiverse Computing, 7 rue de la Croix Martre, 91120 Palaiseau, Paris, France · Department of Basic Sciences, Tecnun – University of Navarra, San Sebastián, Spain · IQM Quantum Computers, Georg-Brauchle-Ring 23-25, 80992 Munich, Germany · Allianz Quantum Hub, Paris, France
Quantum kernel methods are leading candidates for a practical quantum advantage in machine learning, but assessing that potential requires two quantities usually reported separately: how well a kernel performs on the task, and how far its geometry departs from the classical kernels available for the same problem. We introduce a fully unsupervised, multi-objective protocol that optimises simultaneously the normalised pseudo discrepancy (NPD), a label-free proxy for anomaly detection quality, and the geometric difference (GD) to a tuned classical reference kernel, selecting models from the resulting Pareto front. We apply it to malware beaconing detection in real network traffic, using a one-class support vector machine with fidelity and projected quantum kernels over four data encodings, on simulators and on IQM's 20-qubit Garnet processor. NPD-guided selection alone finds a fidelity kernel that beats the tuned classical baseline, but with a geometric difference too small to certify the gain as quantum. Projected kernels reach far larger geometric differences; the Pareto-selected one only marginally exceeds the baseline (AUC 0.782 versus 0.765, gC→Q≈89>N relative to that reference kernel), still below the NPD-selected fidelity kernel (0.840).
Figures & tables
Model
Hyperparameter
Search Space
Projected Quantum Kernel (RBF)
Feature map
Z, IQP, CNOT, Hamiltonian
Number of layers
{1,2,3,4,5}
K-RDM
{1}
Gamma ( γ )
[100μ,10]
ν
[0,1]
Fidelity Kernel
Feature map
Z, IQP, CNOT, Hamiltonian
Table 1: Hyperparameter search space for quantum kernel methods. The description of the considered feature maps is given in Appendix A . The number of layers ranged from 1 to 5. We only considered the single-qubit reduced density matrix for the projected quantum kernels, and γ is the hyperparameter of the projected quantum kernel. ν is a hyperparameter of the OCSVM model, which represents an upper bound on the fraction of training errors and a lower bound of the fraction of support vectors. Curly brackets {} denote discrete values, while square brackets [] denote continuous ranges.
Figure 1: The unsupervised optimisation pipeline: a hyperparameter search over quantum kernels selected under three objectives (NPD, GD, and their joint Pareto knee) and evaluated on the test set.
Feature Name
Forward packets per second
Backward packets per second
Standard deviation of the FFT of forward traffic
Standard deviation of forward inter-arrival time
Standard deviation of packet lengths
Mean forward packet length
Table 2: Feature set obtained from the feature engineering procedure
AUC
F1 (A)
F1 (N)
0.765
0.599
0.561
Table 3: Performance of the classical OCSVM on the malware dataset.
NPD
GD
Pareto
Feature map
Hamiltonian
Hamiltonian
Hamiltonian
Layers
2
1
5
ν
0.322
–
0.266
gC→Q
4.297
7.14
5.94
NPD
2.34
0.592
0.72
Table 4: Best fidelity quantum kernel configuration found for the three objectives on the dataset using the statevector simulator. The GD objective depends only on the kernel matrix, so ν is not defined for that column (–).
NPD
GD
Pareto
Feature map
Z
Hamiltonian
IQP-style
Layers
5
1
5
ν
0.385
–
0.104
γ
2.37
9.99
0.000395
gC→Q
32.0
272
88.6
NPD
2.94
0.603
1.42
Table 5: Best projected quantum kernel configurations found for the three objectives on the six features dataset using the statevector simulator. The GD objective depends only on the kernel matrix, so ν is not defined for that column (–).
Figure 2: Pareto fronts of the fidelity (blue) and projected (red) quantum kernels under joint maximisation of the NPD and GD, evaluated on the statevector simulator. Circles mark the non-dominated configurations of each front; stars mark the selected knee points. The fidelity kernel maintains a low GD across the front, whereas the projected kernel trades a high GD at low NPD for improved NPD, exposing the tension between the two objectives. Each front contains only the configurations that remained non-dominated during the optimisation, which is why the two kernels contribute different numbers of points.
Figure 3: Test AUC versus geometric difference gk for all hyperparameter configurations evaluated during the GD-maximisation run (statevector simulation). Blue: fidelity kernels; red: projected quantum kernels. The green dashed line marks gk=N≈14.14 , below which a classical kernel is guaranteed to match the quantum model; the orange dotted line is the tuned classical OCSVM baseline (AUC =0.765 ). Fidelity kernels stay below the threshold yet reach the highest AUC in the run ( ≈0.84 ), whereas projected kernels extend to gk∼270 with an AUC that shows no systematic trend with gk . A few projected configurations combine gk≫N with an AUC above the baseline, but none reaches the best fidelity kernels, and the two largest-GD models (yellow, green) perform at or below the baseline.
Figure 4: Test AUC versus NPD for all hyperparameter configurations evaluated during the NPD-maximisation run (statevector simulation). Blue: fidelity kernels; red: projected quantum kernels. The orange dotted line marks the tuned classical OCSVM baseline (AUC =0.765 ). Fidelity kernels concentrate at higher AUC, with a substantial fraction above the baseline, while projected kernels spread widely between AUC ≈0.33 and 0.86 and mostly lie below it. Within each family the NPD–AUC relation is weak. The largest-NPD configuration of each family (yellow: fidelity; green: projected) lies above the baseline, at AUC =0.840 and 0.777 respectively.
Method
NPD
GD
Pareto (NPD & GD)
AUC
F1(A)
F1(N)
AUC
F1(A)
F1(N)
AUC
F1(A)
F1(N)
Statevector Simulator
Fidelity
0.840
0.657
0.715
0.666
0.494
0.800
0.812
0.671
0.773
Projected
0.777
0.599
0.563
0.738
0.542
0.340
0.782
0.638
0.736
IQM Noisy Simulator
Fidelity
0.790
0.600
0.582
0.752
0.582
0.630
0.798
0.600
0.576
Table 6: Performance of the optimised quantum kernels obtained from the optimisation of the three objectives: NPD, GD and Pareto frontier using the statevector simulator. We re-run the optimal hyperparameters selected by each optimisation using the IQM noisy simulator; the OCSVM parameter ν was additionally re-tuned on the noisy simulator by maximising the NPD. The classical OCSVM baseline achieves AUC =0.765 . The best value of each metric is shown in bold.
Figure 5: Layout of the IQM Garnet QPU. The color map indicates single and two-qubit gate errors (darker=higher error, lighter=lower error). We highlight the optimal way to choose a qubit patch to implement a Hamiltonian evolution feature map with nearest-neighbour interactions. The specific patch needs to be adjusted depending on re-calibration of the QPU.
Method
Feature Map
AUC
F1(A)
F1(N)
Fidelity
Hamiltonian
0.532
0.640
0.151
Projected
Z
0.612
0.630
0.656
Table 7: Performance of the quantum kernels obtained from the IQM Garnet QPU using the optimal configurations for the hyperparameters as given by the NPD score; for the projected kernel, ν was selected to maximise the test scores.
Similarity in many decision systems is governed not by distance alone but by interactions among variables. In fraud and anomaly detection, small local perturbations can cross interaction-sensitive decision boundaries while leaving ambient distance almost unchanged. Motivated by this setting, we introduce a thin-slab interaction model and an interaction-driven quantum kernel constructed from entangled Pauli-string feature maps. The feature map explicitly encodes sparse high-order block interactions. We show that the resulting fidelity kernel is positive semidefinite, admits an exact block-factorized formulation, and induces a geometry sensitive to changes in interaction regime. Across balanced and imbalanced synthetic experiments spanning third-, fourth-, sixth-, and eighth-order interactions, the proposed kernel consistently outperforms linear, radial basis function, Laplacian, and polynomial kernels, as well as an engineered-interaction linear baseline supplied with the planted block products. On real fraud-detection benchmarks, it achieves the highest mean accuracy and F1 on Credit Card Fraud Detection and ranks second on IEEE-CIS Fraud Detection. Executed on a 156-qubit IBM Quantum processor in a fourth-order setting, the hardware-estimated kernel matches the noise-free simulator within seed-to-seed variability and retains its advantage over the baselines. These findings show that quantum-kernel performance depends on alignment between feature-map geometry and the underlying predictive structure, rather than on Hilbert-space dimension alone. Because the prescribed block-factorized kernel can also be evaluated exactly on a classical computer, the results establish predictive and representational value rather than computational quantum speedup.
Hanqiu Peng, Jianlong Lu, Ying Chen
Centre for Quantitative Finance, Department of Mathematics Risk Management Institute, National University of Singapore, Singapore
A core task in quantum anomaly detection is to compute an anomaly score that quantifies how strongly a test quantum state deviates from a given quantum dataset assumed to be normal. Classically, principal component analysis (PCA) for centered data computes the anomaly score by evaluating the test sample relative to the subspace spanned by the selected leading eigenvectors. However, for quantum data that lack a standard centering, explicitly recovering principal eigenvectors, constructing full Gram matrices, or loading quantum-random-access-memory-style data can be more costly than estimating the anomaly score itself. To avoid these costs, we propose Quantum Spectral Anomaly Detection (QSPADE), which computes PCA-like anomaly scores directly from the spectrum of the average state of the normal dataset. By replacing hard PCA rank selection with a smooth, temperature-controlled spectral threshold, QSPADE makes near-threshold spectral components contribute partially to the anomaly score. This makes the score vary continuously rather than jump when a borderline component is included or excluded, and makes it less sensitive to noise or arbitrary hard cutoffs near the threshold. In the zero-temperature limit, QSPADE recovers the hard-projector PCA score. The proposed measurement-based quantum detector can be calibrated with a sample complexity independent of the data dimension. Numerical simulations show that QSPADE behaves like kernel-PCA on encoded classical data and detects changes across a transverse-field Ising transition without predefined order parameters. Consequently, QSPADE gives an efficient framework for both quantum-kernel anomaly detection on encoded classical data and the monitoring of quantum-native systems where diagnostic observables are unknown.
Yewei Yuan, Michele Minervini, Mark M. Wilde +1
Global College, Shanghai Jiao Tong University, Shanghai 200240, China · School of Electrical and Computer Engineering, Cornell University, Ithaca, New York 14850, United States · Shanghai Jiao Tong University, Shanghai 200240, China +1
Quantum kernel methods have been proposed as a promising approach for leveraging near-term quantum computers for supervised learning, yet rigorous benchmarks against strong classical baselines remain scarce. We present a comprehensive empirical study of quantum kernel support vector machines (QSVMs) across nine binary classification datasets, four quantum feature maps, three classical kernels, and multiple noise models, totalling 970 experiments with strict nested cross-validation. Our analysis spans four phases: (i) statistical significance testing, revealing that none of 29 pairwise quantum-classical comparisons reach significance at α=0.05; (ii) learning curve analysis over six training fractions, showing steeper quantum slopes on six of eight datasets that nonetheless fail to close the gap to the best classical baseline; (iii) hardware validation on IBM ibm_fez (Heron r2), demonstrating kernel fidelity r≥0.976 across six experiments; and (iv) seed sensitivity analysis confirming reproducibility (mean CV 1.4%). A Kruskal-Wallis factorial analysis reveals that dataset choice dominates performance variance (ε2=0.73), while kernel type accounts for only 9%. Spectral analysis offers a mechanistic explanation: current quantum feature maps produce eigenspectra that are either too flat or too concentrated, missing the intermediate profile of the best classical kernel, the radial basis function (RBF). Quantum kernel training (QKT) via kernel-target alignment yields the single competitive result -- balanced accuracy 0.968 on breast cancer -- but with ~2,000x computational overhead. Our findings provide actionable guidelines for quantum kernel research. The complete benchmark suite is publicly available to facilitate reproduction and extension.
Siavash Kakavand, Christoph Strohmeyer, Michael Schlotter
SHARE at FAU, Schaeffler Technologies AG & Co. KG, Herzogenaurach 91074, Bavaria, Germany