quant-phJun 20, 2026

Fine-Tuning Large Language Models for Quantum Reasoning

Authors: Katherine IpCasey R. MyersUdaya ParampalliJames QuachPeiyong Wang

Organizations: School of Computing and Information Systems, The University of Melbourne, Melbourne, Australia · School of Physics, Mathematics and Computing, The University of Western Australia, 35 Stirling Hwy, Crawley WA, 6009, Australia · Pawsey Supercomputing Centre, 1 Bryce Avenue, Kensington WA, 6151, Australia · CSIRO Clayton, Research Way, Clayton VIC 3168, Australia

Abstract

Large language models (LLMs) exhibit abilities beyond natural language modelling and text generation. Recent advances in their reasoning capabilities have spurred interest in applying LLMs to complex scientific tasks requiring deep domain expertise and sophisticated reasoning. Quantum computing, as a highly specialised field with significant knowledge barriers and hardware constraints, could greatly benefit from such advancements. However, a key open question that first must be answered is: How can we develop fine-tuning pipelines that instil genuine quantum reasoning in LLMs, rather than task-specific pattern matching? We study this question through quantum circuit simulation as a training objective, where the model must predict the measurement probability distribution resulting from a sequence of quantum gate operations. We propose and compare two fine-tuning pipelines: (1) Supervised Fine-Tuning (SFT) on explicit gate-by-gate state-vector simulation traces, and (2) a two-stage SFT+Group Relative Policy Optimisation (GRPO) approach that sequentially applies SFT followed by GRPO with verifiable rewards. Our findings show that SFT achieves near-perfect in-distribution and gate-count extrapolation accuracy, significantly outperforming both the base model and the GPT-OSS-120B baseline. SFT+GRPO trades some in-distribution precision for better generalisation to larger qubit systems that SFT alone cannot handle. Both pipelines significantly outperform the baselines, demonstrating that targeted fine-tuning on explicit reasoning traces is an effective strategy for advancing quantum reasoning in LLMs.

Explore similar work

CardsList
  1. Aligning Quantum Operators with Large Language Models

    Jun 11, 2026Rogerio Feris, Yunchao Liu, Pengyuan Li +2Quantum Machine LearningQubit