Quantum Reinforcement Learning (QRL) integrates reinforcement learning with parameterized quantum circuits and is a promising approach to combinatorial optimization. On Noisy Intermediate-Scale Quantum (NISQ) devices, however, decoherence, gate imperfections, and measurement errors reduce policy quality and make learning less reliable. Existing error mitigation techniques are generally applied as fixed corrections that do not adapt to changing noise conditions or to the evolving state of training. This work presents Adaptive Policy-Guided Error Mitigation (APGEM) as a context-aware orchestration layer of the hybrid quantum-classical training loop that dynamically selects the most suitable mitigation strategy during QRL training. APGEM evaluates Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) using policy-level indicators, including quantum-state fidelity, policy entropy, cumulative reward, and approximation ratio, and integrates the selected strategy directly into the reinforcement learning loop. The framework is evaluated on the Capacitated Vehicle Routing Problem (CVRP), a representative NP-hard problem in urban logistics, under a range of NISQ noise models and noise levels. APGEM consistently outperforms conventional static mitigation methods, reaches approximately 94% of the utility of an oracle strategy, maintains higher quantum-state fidelity as noise increases, and produces more stable learning behaviour throughout training. Ablation studies show that the framework learns context-aware mitigation policies that adapt to different noise environments and circuit execution conditions. These findings demonstrate that integrating adaptive error mitigation into the learning process substantially improves the robustness and reliability of QRL on NISQ hardware.
Figures & tables
Ref(s)
Focus of Prior Work
Identified Gap
Contribution of our Work
[ 23 ] , [ 24 ]
Surveys of quantum computing in logistics and supply chains; discuss QAOA, annealing, and hybrid methods.
Do not address noise-resilience or adaptive error mitigation in NISQ settings.
Propose APGEM-QRL , embedding adaptive noise mitigation within a reinforcement learning framework.
[ 25 ] , [ 26 ]
Quantum-inspired algorithms for logistics optimization; scalable heuristics on classical hardware.
Lack of true quantum implementation and no adaptive error management.
Application of QAOA/annealing to logistics problems (TSP, VRP) via QUBO/Ising formulations.
Assumes fixed noise or applies static mitigation; no policy-aware adaptivity.
Introduce policy-guided mitigation that adjusts dynamically to circuit depth, reward dynamics, and noise type.
[ 28 ]
Cognitive/learning-inspired quantum models for route optimization and decision-making.
Conceptual focus; lacks robustness analysis under noisy quantum hardware.
Extend to QRL with adaptive mitigation , ensuring reliable convergence under real-world noise.
[ 29 ] [ 30 ]
Systems perspective on integrating quantum methods into logistics workflows and enterprise KPIs.
Limited demonstrations on realistic urban road networks; no dynamic noise control.
Evaluate on standard CVRPLIB benchmark instances under diverse noise models, reporting best-known-referenced metrics with statistical testing.
[ 26 ]
Data-driven adaptive decision-making in logistics using quantum-inspired techniques.
Adaptivity applied only to classical decision processes, not to quantum error mitigation.
Make error mitigation adaptive , guided by policy feedback and noise diagnostics.
Table 1: Summary of prior works, identified gaps, and contributions of our Work.
Figure 1: Block diagram of the Adaptive Policy-Guided Error Mitigation (APGEM) framework. Quantum circuits generate policies degraded by noise, which are corrected through an adaptive selection of mitigation strategies (ZNE, PEC, CDR, REM). The corrected policy guides QRL updates for solving the CVRP.
Figure 2: Block diagram of the proposed methodology. The workflow begins with VRP instance generation, followed by reinforcement learning environment interaction through a variational quantum circuit. Execution under noise is adaptively corrected by the APGEM controller, with fidelity and entropy metrics guiding policy optimization. The training loop produces results compared against a classical baseline.
Metric
Definition / Formula
Interpretation
Distribution Fidelity ( F )
F(p,q)=(∑ipiqi)2
Classical (Bhattacharyya) fidelity between mitigated and ideal output distributions; higher is better.
Normalized Entropy ( H )
H(p)=−nq1∑ipilog2pi
Shannon entropy of the measured distribution, normalized to [0,1] ; lower indicates a sharper, less noise-dominated distribution.
Cumulative Reward ( R )
R=∑trt
Accumulated shaped reward over an episode; higher corresponds to lower-cost feasible routing.
Routing Cost ( C )
Total travel distance of all routes (with depot returns)
Domain-specific solution quality on the feasibility-guaranteed decode; lower is better.
Approximation Ratio ( AR )
ARBKS=CBKSCquantum (and vs. best classical)
Cost relative to best-known (primary) and to the best classical baseline; values near 1 are near-optimal.
Mitigation Gain ( ΔM )
Mmitigated−Munmitigated
Execution-quality improvement from mitigation at matched conditions.
Table 2: Summary of performance metrics used in evaluating APGEM.
Figure 3: Quantum policy ansatz for the CVRP–QRL framework. The first layer applies Ry(θi) rotations to encode environment-state features into the quantum register. This is followed by a linear entanglement layer of controlled- Z gates to capture correlations across qubits. Two variational layers are then stacked, each consisting of trainable Ry(ϕi(k)) rotations and entangling operations, allowing expressive state transformations. The final measurement distribution guides action selection within the reinforcement learning loop.
Figure 4: Circuit-level illustration of the adaptive error mitigation module integrated into the CVRP–QRL framework.
Component
Version
Python
3.10.11
NumPy
2.2.6
SciPy
1.14.1
pandas
2.2.3
Qiskit
1.4.0
Qiskit Aer
0.16.1
Table 3: Software environment used for all reported results.
Method
Mean AR
95% CI
Avg rank
Feasible
Wilcoxon vs OR-Tools
OR-Tools
1.043
[1.026, 1.062]
1.53
29/30
(reference)
Clarke–Wright
1.048
[1.038, 1.059]
1.77
30/30
p=0.50 ( δ=−0.19 )
Simulated Annealing
1.163
[1.128, 1.206]
2.97
30/30
p=4.6×10−5 ( δ=−0.72 )
Tabu Search
1.276
[1.218, 1.340]
4.10
30/30
p=4.6×10−5 ( δ=−0.82 )
Nearest Neighbour
1.392
[1.351, 1.435]
5.30
30/30
p=3.8×10−6 ( δ=−0.92 )
Genetic Algorithm
1.430
[1.326, 1.550]
5.33
30/30
p=3.8×10−6 ( δ=−0.87 )
Table 4: Classical baselines over 30 CVRPLIB instances. Approximation ratio to best-known (lower is better), 95% confidence interval, average Friedman rank, number of instances solved feasibly, and Wilcoxon signed-rank p -value against the top method (OR-Tools) with Cliff’s delta.
Figure 5: Mean approximation ratio to best-known with bootstrap 95% confidence intervals over 30 CVRPLIB instances. OR-Tools and Clarke–Wright are the strongest and statistically indistinguishable.
Figure 6: Friedman average ranks (1 = best) with the Nemenyi critical difference ( CD=1.38 ). OR-Tools and Clarke–Wright group together and dominate the metaheuristics.
Quantity
Value
Frozen-policy cost (mean)
629.3
Trained-policy cost (mean)
610.8
Improvement over frozen
2.08%
Beats frozen policy
yes
Feasibility rate
1.00
Table 5: Policy learning on P-n16-k8 (three seeds, REINFORCE with parameter-shift gradients). Route cost of the frozen policy versus the trained policy, and feasibility rate.
Figure 7: Route cost versus policy update on P-n16-k8 (mean over three seeds). The REINFORCE learner descends below the frozen-policy baseline; the best classical cost is shown for reference.
Instance
Customers
Qubits
Depth
AR to BKS
Total shots
Runtime/ep (s)
P-n16-k8
15
4
10
1.162
607,232
0.58
A-n32-k5
31
5
12
2.597
946,176
1.14
E-n51-k5
50
6
14
3.132
1,174,016
2.29
M-n101-k10
100
7
16
3.352
2,977,792
14.88
Table 6: Scalability of the quantum policy across instance sizes (contextual bandit controller). Approximation ratio to best-known, qubits, circuit depth, total shots, and mean runtime per episode.
Figure 8: Quantum policy versus classical baselines on A-n32-k5 (best-known =784 , dashed line). Classical solvers, led by OR-Tools at the optimum, remain ahead of the quantum policy at this scale.
Figure 9: Scalability across CVRPLIB instances: approximation ratio to best-known, mean episode runtime, and total shots as the customer count grows from 15 to 100.
Technique
Fidelity
Entropy
CDR
0.835
0.824
PEC
0.829
0.735
ZNE
0.816
0.768
REM
0.769
0.856
No Mitigation
0.711
0.908
Table 7: Per-technique distribution fidelity and normalized entropy on P-n16-k8.
Figure 10: Per-technique distribution fidelity and normalized entropy. Mitigation raises fidelity from 0.711 (unmitigated) to 0.83 (PEC, CDR).
Technique
Shots/ episode
Overhead
Projected HW (s)
No Mitigation
10,240
1.00 ×
81.9
PEC
20,480
2.00 ×
163.8
ZNE
30,720
3.00 ×
245.8
REM
51,200
5.00 ×
409.6
CDR
92,160
9.00 ×
737.3
APGEM (contextual bandit)
40,499
3.96 ×
324.0
Table 8: Sampling overhead per technique on P-n16-k8, relative to unmitigated execution, with per-episode shot counts and a projected cloud-hardware time at 0.4 ms per shot.
Figure 11: Sampling overhead factor per technique relative to unmitigated execution. The adaptive controllers sit between ZNE and REM by mixing techniques according to context.
Noise
No Mitigation
PEC
ZNE
REM
CDR
0.01
0.65
0.09
0.12
0.11
0.03
0.03
0.12
0.22
0.14
0.49
0.03
0.05
0.02
0.74
0.17
0.05
0.01
0.08
0.03
0.19
0.73
0.04
0.02
0.12
0.04
0.15
0.71
0.04
0.06
0.18
0.10
0.09
0.54
0.08
0.18
Table 9: APGEM technique-selection share by depolarizing noise level (2600 decisions). Each technique dominates a different regime.
Figure 12: APGEM technique selection versus depolarizing noise. The dominant technique transitions from no mitigation to readout mitigation, PEC, ZNE and CDR as noise increases.
Figure 13: APGEM technique selection against unmitigated fidelity, entropy and circuit depth, showing that selection is context-dependent.
Figure 14: Windowed selection share over training. All techniques retain nonzero share; the policy does not collapse to a single technique.
Method
Mean reward
Avg rank
Wilcoxon vs APGEM
Oracle
0.639
–
(upper bound)
APGEM
0.602
2.23
(reference)
ZNE
0.585
2.09
p=0.57 ( δ=−0.01 )
PEC
0.569
3.33
p=1.8×10−4 ( δ=0.07 )
CDR
0.522
4.53
p=6.2×10−5 ( δ=0.31 )
REM
0.511
4.69
p=4.3×10−10 ( δ=0.20 )
Table 10: Mean cost-aware reward over 60 held-out conditions and Wilcoxon signed-rank of APGEM against each fixed technique, with Cliff’s delta. The oracle is the expected-utility upper bound.
Figure 15: Mean cost-aware reward across regimes. APGEM exceeds every fixed technique and approaches the expected-utility oracle.
Figure 16: Corrected regret. Left: per-condition expected regret (policy) versus the irreducible noise floor. Right: the raw cumulative regret separated into its noise-floor and expected-regret components.
Figure 17: APGEM reward against the per-condition oracle (left) and cumulative regret (right) over training.
Controller
Mean fidelity
Circuit evals
Best AR
APGEM (Contextual Bandit)
0.820
5550
1.20
APGEM (Bayesian)
0.816
5533
1.16
APGEM (Multi-Objective)
0.799
5214
1.16
Original (Epsilon-greedy)
0.820
5752
1.20
Table 11: APGEM controller comparison on P-n16-k8: mean fidelity, total circuit evaluations, and best approximation ratio.
Noise level
0.01
0.05
0.08
0.10
APGEM
0.966
0.901
0.817
0.794
No mitigation
0.975
0.820
0.712
0.684
Table 12: Mean distribution fidelity versus depolarizing noise level, APGEM against no mitigation (P-n16-k8).
Figure 18: Noise robustness: APGEM sustains higher distribution fidelity than unmitigated execution as noise grows.
Episodes
Best AR
Mean AR
Mean fidelity
Total shots
ZNE share
100
1.200
1.234
0.824
4,659,712
97.5%
500
1.162
1.215
0.823
23,022,080
99.1%
1000
1.162
1.221
0.822
46,104,064
99.5%
Table 13: APGEM metrics versus training horizon on P-n16-k8.
Quantum Reinforcement Learning (QRL) represents policies as variational quantum circuits (VQCs), making it attractive for combinatorial optimization such as the Capacitated Vehicle Routing Problem (CVRP). On noisy intermediate-scale quantum (NISQ) hardware, however, decoherence degrades fidelity and destabilizes learning, and conventional error mitigation is applied statically without regard to the learning context. We introduce Adaptive Policy-Guided Error Mitigation (APGEM), a controller that selects among Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) online, driven by a fidelity, entropy, and cost aware utility function and an epsilon-greedy rule over temporal-difference Q-scores. We evaluate on a realistic urban-logistics testbed, a Delhi-based CVRP over real landmarks with geodesic inter-node costs, exercised across five noise families and four severity levels. On this instance, the QRL agent outperforms constructive heuristics and approaches metaheuristics, while mitigation restores approximation ratios from 0.84-0.87 to 0.92-0.94 under high noise. The controller shifts from a CDR-dominated regime under short training horizons to a balanced deployment across all four techniques under longer horizons, indicating genuine regime-dependent selection. These preliminary results position adaptive, learning-aware mitigation as a practical route to noise-resilient QRL.
Shabir Ahmad Sofi, Bisma Majid, Mir Mohammad Yousuf
Department of Information Technology NIT Srinagar J&K, India
Quantum reinforcement learning (QRL) is a promising approach to learn effective decision strategies across several applications with stochastic environments. Instead of directly modeling the random variables that govern these environments, existing QRL architectures indirectly approximate environment behavior by estimating expected outcomes, which limits their expressive power and adaptive potential. Overcoming such challenges requires a novel QRL approach that exploits the distributional nature of quantum computers to directly model environment random variables as quantum state distributions. Hence, in this paper, a novel framework dubbed quantum-native reinforcement learning (QnRL) is proposed. QnRL is a distributional RL framework that learns conditional distributions naturally in Hilbert space via superimposed and entangled quantum states. Thus, QnRL can directly model the behavior of stochastic learning environments via the natural properties of quantum systems. QnRL accomplishes this via a novel, proposed quantum amplitude kickback (QuAK) algorithm that enables comparing the n-th power of the m-th moment of multiple superimposed distributions. It is theoretically proven that a conditional action policy distribution is distilled from the moments of a quantum generative model entirely within Hilbert space via QuAK, and optimized via QnRL. This complex distribution composition is also shown to provide extra dimensions for expressing environment correlations that are unknown to purely classical and classically-sampled quantum distributional models. Experimental results across diverse environments show that QnRL achieves up to 82.9% higher evaluation scores, with up to 94.3% fewer parameters on average, more accurately estimates the expected return for unseen observations, and better adapts to varying stochastic conditions compared to the baseline.
Alexander DeRieux, Walid Saad
Bradley Department of Electrical and Computer Engineering, Virginia Tech Institute for Advanced Computing, Alexandria, VA 22305 USA
Quantum error mitigation (QEM) is essential for extracting reliable results from near-term quantum devices, yet practical deployments must balance mitigation strength against runtime overhead under time-varying noise. We introduce \emph{GSC-QEMit}, a telemetry-driven, \textbf{context--forecast--bandit} framework for \emph{adaptive} mitigation that switches between lightweight suppression and heavier intervention as drift evolves. GSC-QEMit composes three coupled modules: (G) a Growing Hierarchical Self-Organizing Map (GHSOM) that clusters streaming telemetry into operating contexts; (S) an uncertainty-aware subsampled Gaussian-process forecaster that predicts short-horizon fidelity degradation; and (C) a cost-aware contextual multi-armed bandit (CMAB) that selects mitigation actions via Thompson sampling with explicit intervention cost. We evaluate GSC-QEMit on benchmark circuit families (GHZ, Quantum Fourier Transform, and Grover search) under nonstationary noise regimes simulated in Qiskit Aer, using an instrumented testbed where action labels correspond to graded mitigation intensity. Across Clifford, non-Clifford, and structured workloads, GSC-QEMit improves average logical fidelity by \textbf{+9.0%} relative to unmitigated execution while reducing unnecessary heavy interventions by reserving them for inferred noise spikes. The resulting policies exhibit a favorable fidelity--cost trade-off and transfer across the evaluated workloads without circuit-specific tuning.
Steven Szachara, Sheeraja Rajakrishnan, Dylan Jay Van Allen +3
Department of Software Engineering, Rochester Institute of Technology, Rochester, USA · Institute for Quantum & Information Sciences, Syracuse University, Syracuse, USA