QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG
Organizations: Korea University · Sookmyung Women’s University · Virginia Tech · Seoul National University · Seoul National University Hospital
Abstract
Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent retrieval adaptation (QuPID) repairs this by making the circuit input-dependent through data re-uploading and by comparing measurement readouts, vectors of local Pauli expectations, rather than states. The result is a small readout for adapting frozen image features to a local archive with limited data: training simulates the circuit classically, and inference runs on a GPU with fixed learned parameters. We characterize the class as a structured factorization of input-modulated quadratic feature maps, bound the frequency support of its re-uploading channel, and give a parameter-count generalization bound that motivates its small budget. Under a shared frozen backbone and a label-free protocol, QuPID's 60 parameters give higher precision-at-5 (P@5) on ChestX-ray14 and MURA than frozen medical encoders, and than adapters and low-rank adaptation (LoRA) with up to 5.25 million trainable parameters. On ChestX-ray14, the P@5 gain over the frozen encoder is +0.116, the lead over retuned adapters is widest at 512 adaptation examples (+0.040), and the full-budget margin over an equally compact classical rotation-plane head is +0.023 with a 95% interval excluding zero. Medical imaging is the primary testbed; the pattern recurs on two non-medical benchmarks, in report generation, and under simulated gate noise and finite-shot readout.
Figures & tables
| Method | ChestX-ray14 | MURA | #Params | ||||
| P@5 | MAP@10 | NDCG@10 | P@5 | MAP@10 | NDCG@10 | ||
| ViT Only | 0.312 | 0.278 | 0.307 | 0.481 | 0.451 | 0.475 | 0 |
| BiomedCLIP | 0.352 | 0.315 | 0.348 | 0.508 | 0.476 | 0.501 | 0 |
| MedCLIP | 0.341 | 0.303 | 0.334 | 0.494 | 0.463 | 0.488 | 0 |
| Linear Head | 0.378 .008 | 0.341 .007 | 0.368 .008 | 0.549 .007 | 0.513 .008 | 0.541 .007 | 1.05M |
| Adapter | 0.401 .007 | 0.368 .006 | 0.396 .007 | 0.573 .006 | 0.546 .007 | 0.567 .006 | 525K |
Appendix figures & tables39 assets
Supplementary material from the paper’s appendix.
Appendix
| Work | Venue | Domain | Encoding | Hybrid | Simulator |
| TensorRL-QAS ( Kundu & Mangini, 2025 ) | NeurIPS | RL | Matrix product state | ✓ | ✓ |
| QVF ( Wang et al., 2025d ) | NeurIPS | Vision | Amplitude | ✓ | ✓ |
| QDSFormer ( Born et al., 2025 ) | NeurIPS | Vision | Angle | ✓ | ✓ |
| PQC policies ( Jerbi et al., 2021 ) | NeurIPS | RL | Angle, re-uploading | ✓ | ✓ |
| QuanONet ( Wang et al., 2025c ) | ICML | Operator learning | Angle, re-uploading | ✓ | |
| Quorus ( Han et al., 2026 ) | ICLR | Federated | Angle | ✓ | ✓ |
| Design | vs. frozen | ||
| Naive fidelity design | |||
| QuPID (re-uploading and measurement readout) |
| Backbone | ViT-L/16, ImageNet-21K, frozen (1024-D features) |
| Qubits / depth / readout | , , |
| Trainable parameters | ( rotations, re-uploading scales) |
| Optimizer | AdamW, lr , weight decay |
| Batch size / temperature | 32 / 0.07 |
| Augmentations | rotation , translation , brightness/contrast jitter 0.1/0.1 |
| Epochs / seeds | 50, label-free early stopping (patience 10) / 5 seeds (0 to 4), mean std |
| Domain | Benchmark | 512 examples | Full budget | ||||
| frozen | cls. | QuPID | frozen | cls. | QuPID | ||
| Medical | ChestX-ray14 | 0.312 | 0.347 | 0.401 | 0.312 | 0.414 | 0.428 |
| Industrial | MVTec AD | 0.428 | 0.446 | 0.481 | 0.428 | 0.492 | 0.508 |
| Fine-grained | CUB-200 | 0.612 | 0.628 | 0.658 | 0.612 | 0.671 | 0.684 |
| Variant | P@5 | NDCG@10 | #Trained |
| QuPID (full: , ) | 0.428 .006 | 0.425 .006 | 60 |
| naive fidelity design (Prop. 1 ) | 0.312 .000 | 0.307 .000 | 0 (30 inert) |
| w/o re-uploading, readout kept | 0.412 .008 | 0.409 .008 | 30 |
| readout: single-qubit only ( ) | 0.415 .007 | 0.412 .007 | 60 |
| readout: pair terms only ( ) | 0.402 .008 | 0.398 .008 | 60 |
| depth | 0.402 .007 | 0.399 .007 | 20 |
| Dataset | Retriever | LLaVA-Med (7B) | CheXagent (8B) | ||||||
| B-4 | R-L | MTR | CIDEr | B-4 | R-L | MTR | CIDEr | ||
| ViT Only | 0.134 | 0.298 | 0.167 | 0.241 | 0.147 | 0.312 | 0.178 | 0.263 | |
| Adapter | 0.156 | 0.332 | 0.189 | 0.298 | 0.171 | 0.348 | 0.203 | 0.321 | |
| IU X-Ray | QuPID | 0.163 | 0.341 | 0.194 | 0.294 | 0.178 | 0.357 | 0.208 | 0.318 |
| ViT Only | 0.098 | 0.247 | 0.138 | 0.168 | 0.108 | 0.261 | 0.149 | 0.184 | |
| Adapter | 0.121 | 0.281 | 0.163 | 0.221 | 0.134 | 0.298 | 0.176 | 0.242 | |
| Retriever | LLaVA-Med (7B) | CheXagent (8B) | ||
| CheXbert-F1 | RadGraph-F1 | CheXbert-F1 | RadGraph-F1 | |
| ViT Only | 0.298 | 0.181 | 0.321 | 0.194 |
| Adapter | 0.334 | 0.203 | 0.356 | 0.217 |
| QuPID | 0.352 † | 0.214 † | 0.373 † | 0.229 † |
| Method | P@5 | P@10 | MAP@10 | MRR | NDCG@5 | NDCG@10 |
| ChestX-ray14 | ||||||
| ViT Only | 0.312 | 0.287 | 0.278 | 0.363 | 0.325 | 0.307 |
| Linear Head | 0.378 | 0.351 | 0.341 | 0.434 | 0.389 | 0.368 |
| Adapter | 0.401 | 0.375 | 0.368 | 0.462 | 0.415 | 0.396 |
| Adapter-L | 0.412 | 0.386 | 0.372 | 0.465 | 0.421 | 0.401 |
| Adapter-XL | 0.414 | 0.388 | 0.387 | 0.468 | 0.433 | 0.414 |
| Evaluation | frozen | cls. | QuPID |
| ChestX-ray14 (in-domain) | 0.312 | 0.401 | 0.428 |
| CheXpert (zero-shot) | 0.289 | 0.301 | 0.318 |
| Channel | |||||
| Depolarizing | 0.428 | 0.428 | 0.426 | 0.423 | 0.404 |
| Bit flip | 0.428 | 0.427 | 0.424 | 0.419 | 0.392 |
| Phase flip | 0.428 | 0.428 | 0.427 | 0.426 | 0.418 |
| Amplitude damping | 0.428 | 0.427 | 0.423 | 0.417 | 0.386 |
| Channel | ||||||
| post | noise-aware | post | noise-aware | |||
| Depolarizing | 0.423 | 0.426 | 0.404 | 0.417 | ||
| Bit flip | 0.419 | 0.424 | 0.392 | 0.409 | ||
| Phase flip | 0.426 | 0.427 | 0.418 | 0.423 | ||
| Amplitude damping | 0.417 | 0.423 | 0.386 | 0.405 | ||
| Shots per observable | ||||
| P@5 | 0.389 | 0.418 | 0.426 | 0.428 |
| NDCG@10 | 0.386 | 0.415 | 0.423 | 0.425 |
| Method | ||||||
| QuPID | 60 | 0.031 | 0.016 | 0.008 | 0.004 | 0.002 |
| Quadratic | 164 | 0.051 | 0.027 | 0.014 | 0.007 | 0.004 |
| Adapter | 525K | 0.212 | 0.171 | 0.118 | 0.072 | 0.041 |
| LoRA ( ) | 1.57M | 0.229 | 0.184 | 0.127 | 0.078 | 0.045 |
| sweep baseline | QuPID | margin | |||
| fixed | tuned | fixed | tuned | ||
| 128 | 0.316 | 0.331 | 0.352 | ||
| 512 | 0.347 | 0.361 | 0.401 | ||
| 2,048 | 0.378 | 0.384 | 0.416 | ||
| 8,192 | 0.401 | 0.405 | 0.424 | ||
| 32,768 | 0.414 | 0.416 | 0.428 | ||
| Module | full budget | ||
| Diagonal gain on 60 leading PCA directions | 60 | 0.331 | 0.322 |
| Random Fourier features | 64 | 0.319 | 0.311 |
| Tensor-train head over the ten blocks, bond dimension 2 | 64 | 0.361 | 0.348 |
| Rotation-plane isometry | 60 | 0.405 | 0.386 |
| Quadratic head | 164 | 0.392 | 0.372 |
| Bias-only shift | 1,024 | 0.352 | 0.337 |
| Surrogate | P@5 | in favour of QuPID | |||
| full | at | at full | |||
| Spectrum features, | 640 | 0.318 | 0.341 | ||
| Spectrum features, | 5,120 | 0.339 | 0.372 | ||
| Spectrum features, | 41,000 | 0.357 | 0.396 | ||
| Spectrum features, | 328,000 | 0.362 | 0.407 | [.028,.050] | [.013,.029] |
| Givens rotations, Pauli-type observables | 60 | 0.389 | 0.409 | ||
| Retrieval context | BLEU-4 | CheXbert-F1 | range | |
| IU X-Ray | MIMIC-CXR | MIMIC-CXR | ||
| no retrieval | 0.112 | 0.081 | 0.246 | n/a |
| random retrieval | 0.121 | 0.086 | 0.259 | |
| ViT Only | 0.134 | 0.098 | 0.298 | |
| Adapter | 0.156 | 0.121 | 0.334 | |
| QuPID | 0.163 | 0.127 | 0.352 | |
| Method | site 1 | +site 2 | +site 3 | forget | mean | joint | |
| Adapter | 525K | 0.401 | 0.371 | 0.343 | 0.058 | 0.362 | 0.396 |
| LoRA ( ) | 1.57M | 0.411 | 0.379 | 0.351 | 0.060 | 0.369 | 0.401 |
| QuPID | 60 | 0.428 | 0.416 | 0.402 | 0.026 | 0.398 | 0.412 |
| Method | full budget | epochs | |||
| default | matched | default | matched | ||
| Adapter | 0.343 | 0.351 | 0.401 | 0.404 | 50 80 |
| LoRA ( ) | 0.347 | 0.354 | 0.411 | 0.414 | 50 68 |
| Adapter-XL | 0.338 | 0.349 | 0.414 | 0.417 | 50 61 |
| QuPID | 0.401 | 0.428 | 50 | ||
| QuPID margin | |||||
| Aspect | Interpretation | Evidence |
| Circuit-defined model class | The circuit shares parameters across input-modulated, correlated full-rank quadratic forms. Its contribution is this compact retrieval feature map, supported by competitive ranking quality under classical simulation. | Sec. 1 , 3.3 , 4.2 ; App. A.6 |
| Classical control budgets | The rotation-plane head matches parameters, isometry and output width. Its input-independent rank-one forms differ from the input-modulated full-rank Pauli forms. The comparison evaluates these choices jointly; RFF and the four-projection Quadratic head are more restricted controls. | Table 1 ; App. J.2 , J.3 , M.1 |
| Contribution beyond frozen features | The backbone initialization is shared; post-encoder heads keep features frozen, while LoRA updates attention projections; the untrained circuit and the naive fidelity design are reported as controls, and the naive design exactly matches the frozen baseline as the theory requires; all parameters are updated. | Fig. 4 (a); Table 6 ; App. B.4 , G |
| Finite-scale trainability | Depth , fixed by the feature dimension, local one- and two-body observables; gradient variance is a finite diagnostic, not an asymptotic rate. Classical simulation differentiates the statevector; parameter shift provides the gate-level hardware cost model. | Sec. 3.4 ; Fig. 11 ; Table 13 |
| Simulation cost and scaling | The statevector dimension is , with runtime also scaling in depth, gate count and readout width; per-epoch training time is that of a light adapter, inference time and memory usage are comparable, and the qubit count is swept rather than assumed. | Table 10 ; Fig. 12 (a); App. H.1 |
| GPU deployment and noise simulation | The intended system trains by classical circuit simulation and runs fixed-parameter GPU inference with cached archive readouts. No quantum processor is required, as in quantum parameter-efficient fine-tuning ( Koike-Akino et al., 2025 ; Liu et al., 2025 ) . Noise and shot studies are simulations of an optional hardware realization, not its validation. | Sec. 2 , 4.2 ; Fig. 5 (a,b); App. A.6 , I |
| Aspect | Interpretation | Evidence |
| Scope of the generalization bound | The bound applies to bounded smooth parameter classes and concerns contrastive risk. It does not certify the reported runs or predict P@5 gaps. The empirical claim combines a small budget with competitive retrieval; the gap ratios are separate diagnostics. | Sec. 3.4 ; App. B.3 , J.1 |
| Classical surrogate comparisons | The default rotation-plane head leaves full/512-example margins of / ; tuning reduces them to / . Appendix J.4 extends this to surrogates with far larger budgets and more tuning; without input modulation they stay between and P@5 up to parameters. These measured comparisons do not establish a classical expressivity separation (App. M.1 ). | Sec. 4.2 ; Table 1 , 24 |
| Circuit and readout selection | Real rotations keep the state real and the simulation computationally inexpensive; qubit count, entangling topology, readout width and depth are each swept and saturate at the default, so the operating point is a plateau. | Fig. 4 (a), 12 ; App. A.3 , H.1 |
| Finite-shot readout | A sufficient Hoeffding ranking guarantee is derived with explicit norm-floor, margin and collection-size dependence. Separately, the finite-shot experiment is within P@5 of exact expectations at shots per observable. | Sec. 4.2 ; App. I |
| Dataset coverage | ChestX-ray14 ( images) and MURA for retrieval, IU X-Ray and MIMIC-CXR for generation, MVTec AD and CUB-200 for generality. | Sec. 4.1 ; App. C |
| Low-data accuracy and degradation | The sweep tests retrieval quality as adaptation data shrink. Default rotation-plane margins are at full budget and at 512. The control degrades less; the results support observed accuracy at a small budget, without establishing superior stability. | Sec. 3.4 , 4.1 ; App. J.1 , J.2 |