Organizations: School of Information Engineering, Shanghai Maritime University, Shanghai 201306, China · Research Center of Intelligent Information Processing and Quantum Intelligent Computing, Shanghai 201306, China · Quantum Innovation Centre (Q.InC), Agency for Science, Technology and Research (A*STAR), 2 Fusionopolis Way, Innovis #08-03, Singapore 138634, Republic of Singapore · Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR), 1 Fusionopolis Way, #16-16 Connexis, Singapore 138632, Republic of Singapore · Engineering Cluster, Singapore Institute of Technology, 1 Punggol Coast Road, Singapore 828608, Republic of Singapore · Lee Kong Chian School of Business, Singapore Management University (SMU), Singapore 178899, Republic of Singapore · Fu Foundation School of Engineering and Applied Science, Columbia University, New York, NY 10027, USA · School of Computer Science and Technology, Tongji University, Shanghai 201804, China
Quantum Generative Adversarial Networks (QGANs) have emerged as representative generative models in the Noisy Intermediate-Scale Quantum (NISQ) era and have attracted increasing attention in quantum machine learning. However, most existing QGAN methods rely on patch-based decomposition strategies, which weaken the global consistency of generated images and increase quantum resource overhead. In this work, we investigate a simpler approach: pixel-level, end-to-end image generation using a single-quantum-circuit QGAN. By analyzing the structural matching relationship between the quantum prior and the target data distribution in Hilbert space, we provide a new theoretical perspective for understanding the training behavior of naive end-to-end QGANs. Specifically, we introduce the Quantum Fidelity Landscape (QFL), defined as the pairwise-fidelity structure induced by an ensemble of quantum states and preserved under shared unitary transformations of the quantum generation process. We show that, under a fixed Lipschitz readout, this invariant imposes a one-sided bound on decoded sample separation, motivating calibration of the prior-induced QFL before adversarial training. To validate this theoretical insight, we propose BasicQGAN, a QGAN framework incorporating quantum prior calibration. Before adversarial optimization, BasicQGAN aligns the prior-induced QFL with the data-induced QFL. Experimental results on small-scale grayscale image datasets show that BasicQGAN achieves stable and effective end-to-end pixel-level image generation while requiring fewer qubits and trainable parameters than representative patch-based quantum generators. Furthermore, experiments with different initial quantum-state ensembles show that QFL-calibrated ensembles achieve better generative performance.
Figures & tables
Figure 1 : The green dots in (a) denote a discrete initial-state ensemble sampled from the quantum prior, and the blue dots in (b) represent the corresponding generated-state ensemble after the generator’s unitary evolution. In (b), we also show a denser ensemble example (red dots), whose QFL histogram exhibits a markedly different distribution from the other two ensembles. In each QFL histogram, the horizontal axis denotes the pairwise fidelity between distinct quantum states, while the vertical axis gives the frequency of state pairs falling into each fidelity bin; each unordered pair is counted once, with self-pairs excluded.
Figure 2 : Overall framework of BasicQGAN. (a) During offline QFL calibration, the state-preparation parameters z1,z2,…,zn are optimized to transform the initial ensemble ρ1,ρ2,…,ρn into a calibrated ensemble ρ1′,ρ2′,…,ρn′ . (b) During online adversarial training, the calibrated preparation parameters z1′,z2′,…,zn′ initialize the state ensemble used by the single-circuit quantum generator.
Figure 3 : Empirical QFL distributions computed from 500, 1000, 1500, and 2000 real samples of MNIST digit 0 at 16×16 resolution.
Figure 4 : QFL distributions induced by different prior preparations and by real data. Panels (a)–(c) use uncalibrated sampling schemes, whereas panel (d) uses the QFL-calibrated prior.
Figure 5 : BasicQGAN produces clearer MNIST samples than the end-to-end PQWGAN(Global) baseline at 16×16 resolution.
Figure 6 : BasicQGAN obtains lower FID scores than PQWGAN(Global) across MNIST classes at 16×16 resolution.
Figure 7 : Ablation study on the initial-state preparation algorithm for BasicQGAN on the 16 × 16 digit ’0’ generation task. (a) and (b) show complete training failure resulting from sampling initial parameters from wide intervals, U[0,π] and U[0,1] , respectively. (c) demonstrates severe mode collapse when using an extremely narrow interval, U[0,0.01] . In contrast, (d) shows that our proposed algorithm successfully generates a diverse set of high-quality images.
Model
FID
BasicQGAN with Zi∼U(0,π)
73.93
BasicQGAN with Zi∼U(0,1)
72.64
BasicQGAN Z0∼U(0,0.01) , Zi=0(i>0)
33.97
BasicQGAN (Our algorithm)
20.38
Table 1 : Quantitative comparison of BasicQGAN against variants with different initial-state ensemble preparation methods.
Figure 8 : Visual comparison of generated MNIST samples at 16×16 and 28×28 resolutions. Left: BasicQGAN and WGAN-GP produce relatively smooth and natural images at both resolutions, whereas the images generated by PQWGAN are blurrier and less natural, particularly for digits ‘5’ and ‘9’. Right: magnified BasicQGAN samples at 16×16 resolution (a) and 28×28 resolution (b), where the latter exhibits increased noise artifacts.
Figure 9 : WGAN-GP and BasicQGAN achieve similar FID scores across different classes of the 16 × 16 MNIST dataset, and both consistently obtain lower scores than PQWGAN.
Figure 10 : At 28×28 resolution, BasicQGAN obtains higher FID scores than the baselines, while PQWGAN achieves the lowest FID scores among the compared methods.
Figure 11 : Quantum model performance comparison on the Uppercase Letters and Geometric Shapes datasets. BasicQGAN produces smoother and more natural samples, whereas PQWGAN generates blurrier and less natural outputs.
Figure 12 : BasicQGAN obtains lower FID scores than PQWGAN across different classes of the Uppercase Letter dataset.
Figure 13 : BasicQGAN obtains lower FID scores than PQWGAN across different classes of the Geometric Shape dataset.
Method
Type
MNIST Dataset 16 × 16
MNIST Dataset 28 × 28
Qubit Count
Parameter Count
Qubit Count
Parameter Count
PQWGAN(Global)
Quantum
9
540
11
660
BasicQGAN
Quantum
8
320
10
400
PQWGAN
Quantum
80
800
168
5120
WGAN-GP
Classical
-
945152
-
1732352
Table 2 : Comparison of the resource overhead of the generators from different methods on the MNIST dataset at 16×16 and 28×28 resolutions.
Figure 14 : Comparison between BasicQGAN and BasicQGAN w/ Ancilla qubit on MNIST at 16×16 resolution.
Figure 15 : FID comparison between BasicQGAN and BasicQGAN w/ Ancilla qubit across different MNIST classes at 16×16 resolution.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 16 : QFLs computed from different numbers of real samples for digit 1 in the 16×16 MNIST dataset. Subfigures (a)–(d) use sample sizes of 500, 1000, 1500, and 2000, respectively.
Figure 17 : QFLs computed from different numbers of real samples for digit 2 in the 16×16 MNIST dataset. Subfigures (a)–(d) use sample sizes of 500, 1000, 1500, and 2000, respectively.
Figure 18 : QFL alignment on 16×16 MNIST (digits 0–9). Subfigures (a)–(j) show, for each digit, the QFL of the optimized initial-state ensemble and the true-sample ensemble.
Figure 19 : QFL alignment on 28×28 MNIST (digits 0–9). Subfigures (a)–(j) show, for each digit, the QFL of the optimized initial-state ensemble and the true-sample ensemble.
Figure 20 : Representative samples from our used datasets: (a) MNIST, (b) Uppercase Letters, and (c) Geometric Shapes.
Dataset
Image Resolution
No. of Training Instances
No. of Classes
MNIST
28×28
1500
10 object categories
MNIST
16×16
1500
10 object categories
Uppercase Letters
16×16
1500
10 object categories
Geometric Shapes
16×16
400
4 object categories
Appendix
Table 3 : A summary of the datasets.
Figure 21 : The generated samples for all digit classes (0-9) produced by various models on the 16 × 16 MNIST dataset.
Figure 22 : The generated samples for all digit classes (0-9) produced by various models on the 28 × 28 MNIST dataset.
Figure 23 : The generated samples for all shape classes (angleCross,hexagon, square and straightCross) produced by various models on the 16 × 16 Geometric Shapes dataset.
Figure 24 : The generated samples (A-J) for all letter classes produced by various models on the 16 × 16 Uppercase Letters dataset.
Figure 25 : FID curves over training epochs on MNIST at two resolutions: (a) 16×16 and (b) 28×28 . Each line reports the per-digit FID, showing a consistent decrease and gradual convergence as training proceeds.
Figure 26 : Qualitative comparison of BasicQGAN generated samples on the 16×16 MNIST setting with different generator depths (10, 20, 40, 60, and 80 layers). The 40-layer configuration is used as our default setting.
Layer
10
20
40 (Ours)
60
80
FID
25.12
20.71
20.38
21.74
26.24
Appendix
Table 4 : Comparison of BasicQGAN performance across different generator depths.
Figure 27 : Effect of the different learning rate in Quantum Fidelity Landscape (QFL) optimization algorithm. (a)–(d) histograms of QFL values for the optimized initial-state ensemble and the true-sample ensemble under different learning rates; the corresponding 1D Wasserstein distances are reported in each subplot. (e) Average-fidelity training curves for different learning rates during the optimization, where the dashed line denotes the target average fidelity.
Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circuits. Existing amplitude-based approaches face two key limitations: pixel locations are typically encoded by computational-basis indices or address qubits, causing quantum resources to grow with image resolution; meanwhile, jointly decoding many pixels from normalized quantum states introduces probability competition among pixels and limits precise pixel-wise control. To address these issues, we reformulate quantum image generation as coordinate-conditioned implicit function learning. Our method takes spatial coordinates and latent variables as inputs, uses a classical embedding network to generate input-dependent circuit parameters, and evaluates a variational quantum circuit at each coordinate. Pixel intensities are directly obtained from the expectation value of a dedicated color qubit, and a complete image is generated by querying all spatial coordinates. This design decouples image resolution from address-qubit requirements and avoids shared probability-normalization constraints across pixels. We further design a specialized variational quantum circuit to provide structural inductive bias for coordinate-conditioned generation. Simulated experiments on two benchmark datasets show that our method outperforms FRQI-based generation and PQWGAN in visual and quantitative quality while using fewer qubits, and also achieves better generation quality than the corresponding classical baseline.
Xue Yang, Rigui Zhou, ShiZheng Jia +7
School of Information Engineering, Shanghai Maritime University, Shanghai 201306, China · Research Center of Intelligent Information Processing and Quantum Intelligent Computing, Shanghai 201306, China · Quantum Innovation Centre (Q.InC), Agency for Science, Technology and Research (A*STAR), 2 Fusionopolis Way, Innovis, #08-03, Singapore 138634, Republic of Singapore +5
Quantum generative modeling is a rapidly evolving discipline at the intersection of quantum computing and machine learning. Contemporary quantum machine learning is generally limited to toy examples or heavily restricted datasets with few elements. This is not only due to the current limitations of available quantum hardware but also due to the absence of inductive biases arising from application-agnostic designs. Current quantum solutions must resort to tricks to scale down high-resolution images, such as relying heavily on dimensionality reduction or utilizing multiple quantum models for low-resolution image patches. Building on recent developments in classical image loading to quantum computers, we circumvent these limitations and train quantum Wasserstein GANs on the established classical MNIST and Fashion-MNIST datasets. Using the complete datasets, our system generates full-resolution images across all ten classes and establishes a new state-of-the-art performance with a single end-to-end quantum generator without tricks. As a proof-of-principle, we also demonstrate that our approach can be extended to color images, exemplified on the Street View House Numbers dataset. We analyze how the choice of variational circuit architecture introduces inductive biases, which crucially unlock this performance. Furthermore, enhanced noise input techniques enable highly diverse image generation while maintaining quality. Finally, we show promising results even under quantum shot noise conditions.
Jonas Jäger, Florian J. Kiwit, Carlos A. Riofrío
BMW Group, Munich, Germany · Department of Computer Science and Institute of Applied Mathematics, University of British Columbia (UBC), Vancouver, Canada · Stewart Blusson Quantum Matter Institute (QMI), Vancouver, Canada +1
We propose a quantum implicit neural representation (QINR)-based autoencoder (AE) and variational autoencoder (VAE) for image reconstruction and generation tasks. Our purpose is to demonstrate that the QINR in VAEs and AEs can transform information from the latent space into highly rich, periodic, and high-frequency features. Additionally, we aim to show that the QINR-VAE can be more stable than various quantum generative adversarial network (QGAN) models in image generation because it can address the low diversity problem. Our quantum-classical hybrid models consist of a classical convolutional neural network (CNN) encoder and a quantum-based QINR decoder. We train the QINR-AE/VAE with binary cross-entropy with logits (BCEWithLogits) as the reconstruction loss. For the QINR-VAE, we additionally employ Kullback-Leibler divergence for latent regularization with beta/capacity scheduling to prevent posterior collapse. We introduce learnable angle-scaling in data reuploading to address optimization challenges. We test our models on the MNIST, E-MNIST, and Fashion MNIST datasets to reconstruct and generate images. Our results demonstrate that the QINR structure in VAE can produce a wider variety of images with a small amount of data than various generative models that have been studied. We observe that the generated/reconstructed images from the QINR-VAE/AE are clear with sharp boundaries and details. Overall, we find that the addition of QINR-based quantum layers into the AE/VAE frameworks shows improved performance of reconstruction/generation under the constrained experimental setting relative to the specific baselines.
Saadet Müzehher Eren
Department of Physics, Izmir Institute of Technology, Gülbahçe, Urla, 35430, Izmir, Turkey.