quant-phSep 14, 2026

Towards Surrogate Based Dequantization of Quantum Reinforcement Learning

Authors: Pablo Rodriguez-GrasaSofiene JerbiMikel SanzRyan Sweke

Organizations: Department of Physical Chemistry, University of the Basque Country UPV/EHU, Apartado 644, 48080 Bilbao, Spain · TECNALIA, Basque Research and Technology Alliance (BRTA), 48160 Derio, Spain · Dahlem Center for Complex Quantum Systems, Freie Universität Berlin, 14195 Berlin, Germany · Helmholtz-Zentrum Berlin für Materialien und Energie, 14109 Berlin, Germany · EHU Quantum Center, University of the Basque Country UPV/EHU, Apartado 644, 48080 Bilbao, Spain · IKERBASQUE, Basque Foundation for Science, Plaza Euskadi 5, 48009, Bilbao, Spain · Basque Center for Applied Mathematics (BCAM), Alameda de Mazarredo, 14, 48009 Bilbao, Spain · African Institute for Mathematical Sciences (AIMS), South Africa · Department of Mathematical Sciences, Stellenbosch University, Stellenbosch 7600, South Africa · National Institute for Theoretical and Computational Sciences (NITheCS), South Africa

Abstract

In recent years, the utility of parameterized quantum circuits as function approximators has been widely studied. In the context of reinforcement learning, this approach has led to variational quantum algorithms such as quantum Q-learning. While these methods show promising empirical results, and can provide provable advantages for artificial problems, it remains unclear whether they can provide a provable quantum advantage over classical approaches for problems of practical relevance. A natural way to investigate this question is through the lens of dequantization: The construction of efficient classical algorithms capable of matching the performance of quantum variational methods. Building on recent kernel-based dequantization results for supervised learning, we take steps towards extending this surrogate-based dequantization program to reinforcement learning. Specifically, we study the simplified setting of reinforcement learning with a uniform generative model in which uniformly random state-action samples are available, which models the regime of sampling from a large experience replay buffer after sufficient exploration. Within this setting, we provide finite sample guarantees for classical kernelized Fitted Q-Iteration, with classical kernels designed to match the inductive bias of particular parameterized quantum circuits. Using these results, we then provide a set of sufficient conditions, on the data-encoding strategy of a parameterized quantum circuit, the corresponding classical kernel, and the problem structure, under which kernelized Fitted Q-Iteration provides a meaningful dequantization of quantum Q-learning, in this simplified setting. Apart from providing rigorous dequantization guarantees when these conditions are met, these results also motivate the use of kernelized fitted Q-iteration as a dequantization heuristic when these sufficient conditions cannot be verified.

Explore similar work

CardsList
  1. QnRL: Quantum-Native Reinforcement Learning

    Jun 6, 2026Alexander DeRieux, Walid SaadQ-Learning