cs.LGSep 13, 2026

GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling

Authors: Yufeng WangParivesh PriyeLu WeiHaibin Ling

Abstract

Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport provides fast amortized samples, but reward tuning can reshape the generated distribution rather than improve exploration of this fixed target. We introduce GRPO-QPS, a target-preserving framework in which GRPO learns proposal behavior and an exact Metropolis correction preserves the posterior after training. Across the evaluated reconstruction benchmarks, GRPO-QPS improves over BuresTomFlow and Flow-GRPO on thermal, cat, Dicke, and cluster families, and it closely matches an exact two-qubit reference posterior. Tuned conventional MCMC is slightly stronger on several original continuous benchmarks where the available fixed proposals already match the posterior geometry well. To test whether this reflects a fundamental limitation of learned exploration, we evaluate a more challenging multimodal thermal posterior. At six qubits and 800 shots, the learned proposal achieves a minimum effective sample size of 102 per 1,0001{,}000 likelihood calls, compared with 28 for prior independence, 27 for a tuned fixed mixture, and 20 for Haario adaptive Metropolis. A record-conditioned policy also transfers to unseen 3,000-shot records, matching or exceeding the strongest conventional baseline in all nine held-out seed-record comparisons. These results show that GRPO-QPS combines target-preserving Bayesian inference with broad gains over learned transport baselines and a sampling advantage when efficient exploration requires proposal geometry beyond the evaluated conventional kernels.

Explore similar work

Jul 13, 2026quant-ph

Fixed-Protocol Amortized MPS Tomography with Conformalized Predictive Uncertainty

Quantum state tomography is sample-starved, and the states one prepares live on a narrow, learnable manifold. A k=0k{=}0 prior-only control shows that on concentrated families a prior estimate is already near-optimal, so ``high fidelity at few measurements'' can be family memorization rather than tomography; genuine measurement-efficiency needs a model that conditions on the measurements and demonstrably uses them. On a shared matrix-product-state (MPS) core parameterization we study two routes. ApproachA learns a generative prior over MPS cores with measurement-guided posterior inference (gold-standard-validated, but whose few-measurement accuracy the control shows is largely the prior). ApproachB, our main proposal, is a \emph{fixed-protocol amortized} MPS estimator trained once with a gauge-invariant fidelity loss; we deliberately do not rest it on a permutation-invariant set encoder (a plain MLP matches it). The decisive lever is the measurement design: motivated by the fact that local reduced density matrices determine a χχ-MPS, conditioning on an \emph{informative local} Pauli set rather than random strings turns a modest, memorization-prone estimator into a high-fidelity one ( ⁣0.95\approx\!0.95, up to +0.59+0.59 over prior-only, decisively passing a shuffled-measurement control). A dropout ensemble, conformally recalibrated, gives  ⁣90%\approx\!90\%-coverage intervals -- including for observables never measured, where a shot-based interval does not exist. Quality holds as the system grows (fidelity 0.900.90 at n=10n{=}10, gain \emph{growing} in nn; 0.880.88 at bond dimension χ=4χ{=}4), the parameterization is polynomial (native contraction to 2020 qubits), and we close the loop on IBM hardware (55 states at 0.970.97 from hardware-measured Paulis).
Jian Xu, Delu Zeng, John Paisley +1
Jul 7, 2026cs.AI

QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron

Autonomous systems under partial observability act on beliefs, not raw sensor events. QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event evidence term, and returns an ordinary posterior to a classical planner. This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing posterior. We answer with a controlled hardware case study rather than an end-to-end autonomy or wall-clock speedup claim. The study compares no amplification, guarded Grover amplification, and all-step fixed-point amplification on the same trajectory, then checks whether the returned posterior would change the downstream action. All-step FPAA preserves the Tiger posterior across the reported 8-step and 12-step primary runs, and the 20-step and 32-step controls remain inside the same operating band. In every reported decision check, the hardware posterior and the exact Bayes posterior select the same immediate action. Boundary-aware BIQAE stabilizes amplitude estimation near zero and near one, while a rare-event sweep maps the logical sample-complexity envelope for one-in-a-million evidence. The result is an operating envelope for a hardware-calibrated belief-update primitive, not a standalone hardware-advantage claim.
Bayram Yuksel Eker, Suayb S. Arslan, Ozgur Nazli +2
Sep 12, 2026quant-ph

Generative Replay Mitigates Sample Starvation in Quantum Architecture Search

Reinforcement learning (RL) can automate quantum architecture search, but its scalability is limited when useful circuit trajectories become rare in the rapidly expanding search space. Existing replay mechanisms reuse observed transitions; the proposed learned model produces additional predicted one step transitions from real state-action seeds. Here we introduce GenQAS, a tensor network-guided RL framework that combines a fixed matrix product state warm-start with prioritized generative replay. A learned local transition model generates synthetic circuit transitions on demand and mixes them with real experience during Double Deep Q-Network updates. Under a random exploration analysis, near ground state circuits occupy a rapidly shrinking region of the accessible state space. We investigate whether real data anchored synthetic replay can improve the effective training signal in this regime. Across chemical Hamiltonian benchmarks from 6 to 12 qubits, GenQAS improves fixed-budget success probability and identifies compact circuits at competitive energy error. At 12 qubits, it improves final success probability by up to 7.0×7.0\times over passive replay. On a 15-qubit transverse field Ising model, GenQAS increases success probability from 12%12\% to 21%21\%. In a noisy 6-qubit BeH2_2 transfer experiment, generative replay reduces the steps to chemical accuracy by 92.7%92.7\%. These results show that generative replay can mitigate sample starvation in quantum architecture search and support more resource efficient circuit discovery.
Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld +1