Organizations: Nanyang Technological University, Singapore · Hong Kong University of Science and Technology, Hong Kong, China · Zhejiang University of Technology, Hangzhou, China
Quantum libraries are now critical infrastructure for quantum algorithm development, yet their correctness remains difficult to test. Existing testing techniques mainly rely on failure-based or comparison-based oracles, exposing bugs only when executions fail, violate runtime checks, or disagree with another implementation. Their applicability is limited when suitable execution-based oracles are unavailable, leaving some silent bugs undetected. Such missed bugs can produce incorrect results that propagate into experimental conclusions, simulation studies, and algorithmic designs. Here we present QuSema, an autonomous testing agent for finding silent bugs in quantum libraries. QuSema uses constraints from quantum semantics and documentation as a source-level semantic oracle to assess whether implementation logic can produce invalid outputs from valid inputs. It operates through an agentic loop that repeatedly inspects library API documentation and source code, reasons about the intended behavior of quantum operations, identifies potential semantic deviations, and validates them by generating executable tests through library APIs. Guided by quantum-domain reasoning, QuSema turns high-level behavioral mismatches into concrete, user-triggerable bug reports, enabling it to uncover non-crash defects. We implement QuSema for Qiskit and PennyLane. On a benchmark of 20 historical silent bugs, QuSema achieves higher mean bug relocation counts than Claude Code and Codex, with the DeepSeek configuration costing less than Claude Code. QuSema also discovers 40 previously unknown bugs confirmed by the developers, including 30 silent bugs.
Figures & tables
Figure 1. Limitations of approaches employing black-box test oracles. (a) Limited availability of expected relations between execution results, where equivalent circuits can have different resource counts and some higher-level resource-estimation functionality lacks a cross-library counterpart. (b) Inefficient triggering, where a bug when implementing the CU gate is exposed only under specific semantic conditions.
Figure 2. Circuit-level manifestation and source-level semantic evidence of Qiskit #13079 ( Qiskit Developers, 2024b ) in HoareOptimizer . (a) The triggering program and actual and expected circuits, showing an extra X on q0 . (b) The implementation caches the original node after replacing it, so subsequent cancellation leaves the replacement X in the circuit. This violates quantum circuit semantics and is inconsistent with the documented node replacement.
Figure 3. Workflow of QuSema for detecting silent bugs in quantum libraries. (a) Contextual Unit Preparation extracts code segments from library source code and combines them with call relations and documentation to form contextual units. (b) Semantic Defect Detection searches each contextual unit for semantic violations, formulates defect hypotheses and executes tests to produce candidate defects. (c) API Triggering and Bug Validation searches for documented APIs that expose the candidate defects, uses the reverse chain as auxiliary guidance, verifies the triggers, and classifies the results as bugs , potential bugs , or not bugs .
Category
Functionality type
Library
Issue
Version
Comparable functionality
Quantum programming language
Semantic queries
PennyLane
#4842 ( PennyLane Developers, 2023 )
0.33.1
Unavailable
Semantic queries
Qiskit
#14107 ( Qiskit Developers, 2025b )
2.2.0
Unavailable
Construction
PennyLane
#6359 ( PennyLane Developers, 2024 )
0.38.0
Available
Construction
Qiskit
#15465 ( Qiskit Developers, 2025a )
2.4.1
Available
Specialized quantum algorithms
PennyLane
#9521 ( PennyLane Developers, 2026w )
0.45.0
Available
Specialized quantum algorithms
Qiskit
#12093 ( Qiskit Developers, 2024d )
1.0.2
Available
Table 1. Historical silent bug benchmark covering 20 maintainer-confirmed bugs in Qiskit and PennyLane. Comparable functionality indicates whether cross-library comparison is available.
Library
Version
LoC
APIs
Code segments
Min. lines
Max. lines
Qiskit
2.4.1
169,750
899
8,949
1
908
PennyLane
0.45.0
104,172
1,270
7,786
2
446
Table 2. Library statistics for RQ3. APIs count entries listed in the official documentation. Library LoC excludes comments and blank lines, while segment lengths include them and exclude attached context.
Configuration
Run 1
Run 2
Run 3
Mean
Costs (USD)
QuSema + Opus 5
16
15
17
16.00
1,603.44
QuSema + DeepSeek
15
16
16
15.67
162.27
Claude Code + Fable 5
15
14
15
14.67
425.04
Codex + GPT-5.6-Sol
15
12
13
13.33
117.29
Table 3. Bug relocation and estimated cost on the historical benchmark. Runs 1–3 report manually reviewed counts, Mean is their average, and Cost covers all three runs.
Configuration
Run
Candidates
Verified defects
API triggers
Final reports
QuSema + Opus 5
1
65
64
64
64
2
56
55
54
54
3
66
66
66
66
QuSema + DeepSeek
1
64
61
60
60
2
73
72
72
72
3
65
62
62
62
Table 4. Candidate outcomes for all findings produced during historical bug relocation after deduplication.
Figure 4. Component contributions with DeepSeek-V4.1-Flash (Max). (a) Candidates matching historical defects as Contextual Unit Preparation and the Source-level Semantic Oracle are added. (b) Bugs satisfying the relocation criterion with and without hypothesis verification in Stages II and III. Each group reports the counts from Runs 1–3 and their mean across the three runs. Two bar charts on the same row. Each group contains four bars for Runs 1–3 and their mean. The left chart shows identified historical defects: Base, 12, 14, and 15, with mean 13.67; adding Contextual Unit Preparation, 15, 15, and 15, with mean 15.00; further adding the Source-level Semantic Oracle, 15, 16, and 16, with mean 15.67. The right chart shows relocated historical bugs: without hypothesis verification, 14, 16, and 16, with mean 15.33; full QuSema, 15, 16, and 16, with mean 15.67.
Figure 5. Silent bugs discovered by QuSema: (a) incorrect bosonic-operator multiplication in PennyLane #9650 ( PennyLane Developers, 2026b ) and (b) incorrect control-state composition in Qiskit #16410 ( Qiskit Developers, 2026n ) . Two Python reproducers. In panel (a), BoseWord objects constructed from identical dictionary entries in different insertion orders compare equal, but multiplication returns different products. In panel (b), OptimizeAnnotated combines nested controls incorrectly and changes the circuit unitary.
As quantum computing continually improves, ensuring the reliability and correctness of quantum libraries has become increasingly critical. To this end, many LLM-based fuzzing approaches towards quantum libraries have been proposed to uncover potential bugs. However, these methods still suffer from limitations such as insufficient flexibility and low efficiency, which hinder the progress of the quantum computing field. To address these challenges, we propose KQFuzz, a novel knowledge-guided fuzzer for quantum libraries. It leverages comprehensive codebase knowledge to ground LLM-based test generation, synergizing this with fitness-guided evaluation and two-level mutations to explore complex execution paths and trigger potential bugs. Firstly, KQFuzz introduces a novel prompting scheme tailored to quantum programs, which strategically incorporates knowledge of the codebase to efficiently generate high-quality quantum seed programs. Moreover, we develop evaluation and mutation strategies to handle the generated seed programs, facilitating efficient fuzzing execution while further enriching the diversity of the resulting test cases. We implement KQFuzz and conduct fuzzing on three popular quantum libraries, including Qiskit, PennyLane, and Cirq. Experimental results demonstrate that our approach significantly outperforms other state-of-the-art methods, with coverage improved by up to 18.44%. During the development of KQFuzz, we discovered 13 bugs, all of which have been confirmed and 12 have already been fixed by the developers.
Fuyuan Xia, Qixin Zhang, Chenhao Ying +5
1Shanghai Jiao Tong University · 2Nanyang Technological University · 3Hong Kong University of Science and Technology +1
Quantum machine learning integrates the strengths of quantum computing and machine learning, enabling models to learn complex features using fewer parameters than their classical counterparts. Due to the increasing complexity of quantum machine learning models, it is necessary to verify that the implementation of these models satisfy the design specification and be free of bugs and faults. Mutation testing is a promising avenue to identify faulty quantum circuits that do not meet design specifications or contain defects by intentionally inserting faults into the quantum circuit. It is necessary to define mutation operations to inject faults into quantum circuits to ensure that a test suite is robust enough to evaluate an implementation against its design specification. In this paper, we extend mutation testing to quantum machine learning applications, primarily quantum neural network models. Specifically, this paper makes two important contributions. We define new mutation operations for efficient fault insertion compared to state-of-the-art approaches. We also present a directed mutation generation technique to reduce redundant mutant circuits. Extensive experimental evaluation demonstrates that our approach generates a more diverse and representative set of mutants, effectively addressing faults that traditional techniques fail to expose.
We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing framework, and package it with an evaluation framework for systematic LLM assessment. The resulting benchmark comprises 350 tasks across 26 categories, spanning fundamental gates through advanced algorithms (Grover's, Simon's, Deutsch-Jozsa), error correction, key distribution, and quantum games. Each task includes a natural language prompt, canonical solution, and deterministic test verification via classical circuit simulation. By building on the QuantumKatas' proven pedagogical design rather than creating tasks from scratch, we inherit a principled difficulty progression and comprehensive concept coverage while contributing the framework adaptation, evaluation infrastructure, and empirical analysis. We evaluate 16 LLMs across 7 prompting configurations -- a total of 39,200 model runs -- to demonstrate the benchmark's utility. Three key findings emerge: (1) the benchmark effectively differentiates model capabilities, with best-configuration pass rates ranging from 32.3% to 83.1% and a 26.1 pp average gap between frontier and open-source models; (2) models perform well at implementing known algorithms (SimonsAlgorithm 82.1%, BasicGates 81.6%) but struggle with problem encoding (SolveSATWithGrover 34.4%, DistinguishUnitaries 40.0%); and (3) chain-of-thought prompting shows a modestly bimodal effect -- it is the best strategy for three models (two of them explicitly reasoning-tuned per vendor documentation) but degrades performance for the rest, leaving it mid-pack in aggregate (56.3% mean) behind few-shot-5 (57.8%). We release the benchmark, evaluation framework, and baseline results to support research on LLM capabilities in quantum computing.