SQUARE: Structured Quantum Representation Adapters as Compact Quadratic Feature Maps for Frozen Language Models
Organizations: Korea University · Sookmyung Women’s University · Ulsan National Institute of Science and Technology · Purdue University · Seoul National University Hospital
Abstract
Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain linear at the adaptation module itself, whereas explicit second-order alternatives introduce pairwise interactions through direct parameterization or predefined factorizations. We propose SQUARE, a Structured QUAntum REpresentation adapter that amplitude-encodes the bottleneck vector, applies a parameterized quantum circuit, and measures the resulting state. We show that each basis-probability feature is exactly a normalized quadratic form in the bottleneck coordinates, while the additional Pauli- readouts are signed linear combinations of these probabilities. The measured map can therefore parameterize interactions over coordinate pairs through a small set of shared circuit parameters, where is the bottleneck dimension. It provides a structured parameterization within, rather than beyond, the classical normalized-quadratic feature class. In a disjoint same-pipeline evaluation over eight GLUE-derived controlled interaction tasks and five shared seeds, SQUARE achieves an average test accuracy of , compared with for an affine normalized-quadratic predictor, for the evaluated parameter-matched Givens mixing model, for an MLP, and for a frozen-circuit control. Under reduced supervision, it also shows consistent gains over the strongest evaluated classical comparator, with the same qualitative pattern across multiple frozen LM backbones. All circuit experiments use simulation, while the learned feature map can be evaluated exactly in batched PyTorch without quantum hardware.
Figures & tables
| Method | Params | Acc. | F1 | AUC |
| BitFit | 33 | 0.6502 | 0.6463 | 0.6641 |
| LoRA r=4 | 104 | 0.6383 | 0.6283 | 0.6536 |
| LoRA r=8 | 200 | 0.6444 | 0.6378 | 0.6576 |
| Prefix | 289 | 0.6595 | 0.6562 | 0.6734 |
| MLP | 776 | 0.6869 | 0.6707 | 0.7217 |
| SQUARE (Ours) | 180 | 0.7416 | 0.7536 | 0.8259 |
| Method | Params | Boundary | 10% | 25% | 50% | 100% |
| Classical Fourier | 210 | 0.980 | 0.736 0.125 | 0.830 0.160 | 0.879 0.156 | 0.981 0.020 |
| Explicit polynomial | 386 | 0.975 | 0.688 0.049 | 0.851 0.083 | 0.820 0.167 | 0.969 0.019 |
| MLP | 961 | 0.598 | 0.515 0.026 | 0.599 0.027 | 0.563 0.092 | 0.616 0.122 |
| SQUARE (Ours) | 281 | 0.992 | 0.823 0.098 | 0.911 0.052 | 0.948 0.026 | 0.979 0.016 |
| Comparator | Gap | Primary contrast |
| Frozen circuit | Trained vs. fixed mixing | |
| Head only | Measured vs. no measured map | |
| MLP | Interaction-feature architecture | |
| Givens-12 | Matched-angle mixing family | |
| Norm. quadratic | Coupled vs. affine quadratic score | |
| SQUARE avg. test accuracy: 0.7565 | ||
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Coefficient-matrix structure | Stored | Effective family |
| Full explicit quadratic | any symmetric | ||
| Low-rank bilinear (rank ) | rank ; dim. | ||
| Factorization machine (rank ) | off-diagonal coefficients from | Gram-constrained off diagonal | |
| Learned mixing + squares | |||
| Learned orthogonal mixing + squares | spectral form, | ||
| SQUARE, real ( ) | restricted orthogonal squares | sampled local rank before |
| Characteristic | QAA | QPA | SQUARE (Ours) |
| Purpose | Adapt states | Generate weights | Relation modeling |
| Data encoding | Amplitude | None | Amplitude |
| Hybrid QC | ✔ | ✔ | ✔ |
| Quantum simulator | ✔ | ✔ | ✔ |
| Experiment | NLG | Perplexity eval. | Controlled / Reranking |
| Work | Conf | Category | Hybrid Q-C | Encoding | Simulator |
| TensorRL-QAS ( Kundu and Mangini, 2025 ) | NeurIPS | RL | ✔ | Matrix Product State | ✔ |
| QVF ( Wang et al., 2025c ) | NeurIPS | CV | ✔ | Amplitude | ✔ |
| AQC-DRC ( Xiao et al., 2025 ) | NeurIPS | ML | ✔ | Amplitude | ✔ |
| QDSFormer ( Born et al., 2025 ) | NeurIPS | CV | ✔ | Angle | ✔ |
| QuanONet ( Wang et al., 2025b ) | ICML | ML | ✔ | Feature Map | ✔ |
| QRL ( Meyer et al., 2025 ) | ICML | RL | ✔ | Angle/Re-uploading | ✔ |
| Method | Params | Checkerboard | Radial ring | Average |
| Classical Fourier | 210 | 0.977 0.033 | 0.983 0.004 | 0.980 |
| Explicit polynomial | 386 | 0.973 0.001 | 0.977 0.003 | 0.975 |
| MLP (GELU) | 961 | 0.487 0.009 | 0.710 0.041 | 0.598 |
| Adapter transformation | 140 | 0.505 0.071 | 0.638 0.012 | 0.571 |
| LoRA transformation | 77 | 0.567 0.025 | 0.517 0.009 | 0.542 |
| SQUARE (Ours) | 281 | 0.987 0.009 | 0.997 0.002 | 0.992 |
| Method | Params | SST-2 | RTE | MRPC | QNLI | Avg. |
| MLP ∗ | 577 | 0.964 0.002 | 0.976 0.002 | 0.958 0.002 | 0.956 0.003 | 0.964 |
| Explicit norm. quadratic ∗ | 409 | 0.799 0.006 | 0.859 0.002 | 0.854 0.004 | 0.827 0.005 | 0.835 |
| Full bilinear ∗ | 273 | 0.874 0.027 | 0.918 0.019 | 0.902 0.011 | 0.908 0.018 | 0.901 |
| Low-rank bilinear | 81 | 0.971 0.005 | 0.967 0.005 | 0.942 0.011 | 0.962 0.009 | 0.961 |
| Factorization machine | 71 | 0.973 0.002 | 0.972 0.002 | 0.965 0.005 | 0.963 0.003 | 0.968 |
| SQUARE (Ours) | 68 | 0.967 0.008 | 0.980 0.005 | 0.972 0.002 | 0.958 0.009 | 0.969 |
| Protocol | Method | Params | SST-2 | RTE | MRPC | QNLI | Avg. |
| Encoder-adapted | LoRA | 221,953 | 0.870 | 0.581 | 0.832 | 0.736 | 0.755 |
| Frozen bottleneck | MLP | 5,041 | 0.822 | 0.540 | 0.811 | 0.571 | 0.686 |
| Frozen bottleneck | SQUARE (Ours) | 49 | 0.826 | 0.527 | 0.815 | 0.578 | 0.686 |
| Frozen bottleneck | Explicit quadratic | 445 | 0.824 | 0.475 | 0.813 | 0.580 | 0.673 |
| Method | Reported learned scalars | Test AUC |
| SQUARE (Ours) | 234 | 0.984 0.014 |
| SVM (RBF) | — | 0.974 0.030 |
| Classical Fourier | 204 | 0.961 0.050 |
| Graph-Laplacian | — | 0.953 0.003 |
| Explicit polynomial | 334 | 0.947 0.030 |
| kNN ( ) | — | 0.909 0.011 |
| Backbone | MLP | SQUARE | Gain | MLP params | SQUARE params |
| OPT-350M | 0.7960 | 0.8294 | 776 | 180 | |
| GPT-2 | 0.7847 | 0.8258 | 776 | 180 | |
| OpenLLaMA-3B | 0.7052 | 0.7714 | 1,633 | 157 | |
| Mistral-7B | 0.7012 | 0.7481 | 1,633 | 157 |
| Method | Quantum | Params | CoLA | SST-2 | STS-B | QQP | MNLI | QNLI | RTE | WNLI | Avg. |
| BitFit | ✗ | 33 | 0.4775 | 0.4825 | 0.4325 | 0.4700 | 0.5125 | 0.4600 | 0.5379 | 0.5070 | 0.4850 |
| LoRA r=4 | ✗ | 104 | 0.6000 | 0.6050 | 0.6450 | 0.6700 | 0.7225 | 0.7675 | 0.8159 | 0.7887 | 0.7018 |
| LoRA r=8 | ✗ | 200 | 0.6500 | 0.6125 | 0.6460 | 0.6700 | 0.7275 | 0.7679 | 0.8231 | 0.7928 | 0.7112 |
| AdaLoRA | ✗ | 621 | 0.6850 | 0.6025 | 0.6450 | 0.6675 | 0.7200 | 0.7675 | 0.8303 | 0.7324 | 0.7063 |
| Prefix | ✗ | 289 | 0.4775 | 0.4825 | 0.4325 | 0.4700 | 0.5125 | 0.4600 | 0.5379 | 0.5070 | 0.4850 |
| MLP | ✗ | 776 | 0.7775 | 0.8025 | 0.7925 | 0.7950 | 0.7825 | 0.7775 | 0.8520 | 0.7887 | 0.7960 |
| Method | Quantum | Params | CoLA | SST-2 | STS-B | QQP | MNLI | QNLI | RTE | WNLI | Avg. |
| BitFit | ✗ | 33 | 0.4825 | 0.4900 | 0.5075 | 0.4425 | 0.4375 | 0.4525 | 0.4946 | 0.5070 | 0.4768 |
| LoRA r=4 | ✗ | 104 | 0.7700 | 0.7425 | 0.7800 | 0.6300 | 0.7525 | 0.7325 | 0.6029 | 0.7528 | 0.7204 |
| LoRA r=8 | ✗ | 200 | 0.7775 | 0.7475 | 0.7875 | 0.6675 | 0.7550 | 0.7325 | 0.6282 | 0.7810 | 0.7346 |
| AdaLoRA | ✗ | 621 | 0.7750 | 0.7575 | 0.7825 | 0.6375 | 0.7550 | 0.7325 | 0.6534 | 0.8028 | 0.7370 |
| Prefix | ✗ | 289 | 0.4850 | 0.4900 | 0.5075 | 0.4425 | 0.4375 | 0.4525 | 0.4946 | 0.5070 | 0.4771 |
| MLP | ✗ | 776 | 0.7625 | 0.7925 | 0.8300 | 0.6950 | 0.8275 | 0.7675 | 0.7978 | 0.8046 | 0.7847 |
| Method | Quantum | Params | CoLA | SST-2 | STS-B | QQP | MNLI | QNLI | RTE | WNLI | Avg. |
| Linear | ✗ | 17 | 0.6358 | 0.5783 | 0.6083 | 0.6583 | 0.6083 | 0.6392 | 0.6474 | 0.6150 | 0.6238 |
| BitFit | ✗ | 33 | 0.4808 | 0.5375 | 0.5592 | 0.5158 | 0.4717 | 0.4908 | 0.5535 | 0.5352 | 0.5181 |
| LoRA r=8 | ✗ | 417 | 0.6300 | 0.5825 | 0.6058 | 0.6525 | 0.6075 | 0.6417 | 0.6450 | 0.6103 | 0.6219 |
| AdaLoRA | ✗ | 621 | 0.6292 | 0.5725 | 0.5892 | 0.6592 | 0.6050 | 0.6400 | 0.6510 | 0.6103 | 0.6195 |
| MLP | ✗ | 1633 | 0.6667 | 0.6425 | 0.7442 | 0.6858 | 0.6767 | 0.6367 | 0.6967 | 0.8920 | 0.7052 |
| SQUARE (Ours) | ✔ | 157 | 0.7667 | 0.7358 | 0.7633 | 0.7825 | 0.7833 | 0.8025 | 0.8231 | 0.7136 | 0.7714 |
| Method | Quantum | Params | CoLA | SST-2 | STS-B | QQP | MNLI | QNLI | RTE | WNLI | Avg. |
| Linear | ✗ | 17 | 0.6017 | 0.5692 | 0.6167 | 0.6225 | 0.6083 | 0.5758 | 0.5860 | 0.6244 | 0.6006 |
| BitFit | ✗ | 33 | 0.4950 | 0.4883 | 0.5200 | 0.5517 | 0.4842 | 0.4500 | 0.5090 | 0.5634 | 0.5077 |
| LoRA r=8 | ✗ | 417 | 0.5975 | 0.5675 | 0.5800 | 0.6308 | 0.6133 | 0.5783 | 0.5752 | 0.6291 | 0.5965 |
| AdaLoRA | ✗ | 621 | 0.5958 | 0.5700 | 0.5808 | 0.6233 | 0.6000 | 0.6000 | 0.5909 | 0.5962 | 0.5946 |
| MLP | ✗ | 1633 | 0.6225 | 0.6592 | 0.7625 | 0.6650 | 0.6567 | 0.6475 | 0.6811 | 0.9155 | 0.7012 |
| SQUARE (Ours) | ✔ | 157 | 0.7442 | 0.7442 | 0.7100 | 0.7158 | 0.7850 | 0.7642 | 0.7377 | 0.7840 | 0.7481 |
| Backbone | Reranker | Label Acc. | Label F1 | Cand. AUC |
| OpenLLaMA-3B | Random | 0.485 | 0.490 | – |
| Linear | 0.515 | 0.679 | 0.522 | |
| MLP | 0.707 | 0.701 | 0.736 | |
| SQUARE (Quantum) | 0.704 | 0.712 | 0.745 | |
| SQUARE (Hybrid) | 0.708 | 0.724 | 0.745 | |
| Mistral-7B-v0.1 | Random | 0.510 | 0.491 | – |
| Method | Params | Fwd. ms | Train ms | Train samples/s |
| BitFit | 33 | 0.0036 | 0.0403 | 24812.7 |
| LoRA- | 104 | 0.0060 | 0.0467 | 21430.0 |
| LoRA- | 200 | 0.0059 | 0.0462 | 21627.3 |
| Prefix | 289 | 0.0112 | 0.0660 | 15153.3 |
| MLP | 776 | 0.0058 | 0.0513 | 19477.8 |
| QAA | 140 | 8.9023 | 14.7052 | 68.0 |
| Implementation | ms / step | ms / sample | relative to MLP |
| SQUARE, PennyLane default.qubit + backprop | |||
| SQUARE, native PyTorch statevector | |||
| MLP ( parameters) |
| Selector | BLEU | ROUGE-1 | ROUGE-2 | ROUGE-L | Score | HitRate | NonlinearScore |
| Uniform random (exact) | 0.0193 | 0.2078 | 0.0622 | 0.1749 | 0.2066 | 0.1656 | 0.4991 |
| MLP | 0.0215 | 0.2256 | 0.0664 | 0.1861 | 0.2149 | 0.2083 | 0.4917 |
| SQUARE, pure path | 0.0211 | 0.2172 | 0.0605 | 0.1775 | 0.2128 | 0.1917 | 0.5222 |
| SQUARE, residual path | 0.0246 | 0.2365 | 0.0757 | 0.1895 | 0.2224 | 0.2417 | 0.5032 |
| Oracle- | 0.0288 | 0.3123 | 0.1098 | 0.2622 | 0.2861 | 1.0000 | 0.5662 |
| Oracle- | 0.0172 | 0.1935 | 0.0604 | 0.1618 | 0.2261 | 0.3083 | 0.6890 |
| Method | cluster | radial ring | checkerboard | two-moons | islands |
| SQUARE (depth-1, 93p) | 0.999 0.001 | 0.825 0.045 | 0.849 0.039 | 0.964 0.034 | 0.986 0.013 |
| SQUARE untr. ( 81p) | 0.999 0.001 | 0.841 0.034 | 0.838 0.042 | 0.981 0.016 | 0.964 0.023 |
| SQUARE (depth-2, 121p) | 0.999 0.001 | 0.999 0.001 | 0.962 0.033 | 0.999 0.001 | 0.994 0.004 |
| SQUARE untr. | 0.999 0.002 | 0.933 0.043 | 0.829 0.015 | 0.997 0.003 | 0.988 0.007 |
| MLP ( 129p) | 0.999 0.001 | 0.804 0.106 | 0.484 0.046 | 0.920 0.004 | 0.992 0.009 |
| Adapter ( 125p) | 0.999 0.002 | 0.773 0.110 | 0.490 0.040 | 0.920 0.003 | 0.994 0.007 |
| Method (rule) | exact | 100 shots | 1,000 | 10,000 |
| SQUARE (checkerboard) | 0.849 | 0.833 | 0.849 | 0.849 |
| SQUARE (checkerboard) | 0.962 | 0.954 | 0.962 | 0.963 |
| SQUARE (radial ring) | 0.999 | 0.997 | 0.999 | 0.999 |
| SQUARE (two-moons) | 0.999 | 0.999 | 0.999 | 0.999 |
| Protocol | Head architecture | Angles | Proj. | Head | Total |
| V20: validation/backbone suites and runtime benchmark | – – , bias-free, on the 20-dimensional readout | 12 | – | 168 | 180 |
| H16: primary held-out test and post-entanglement ablation (Apps. B.12 , B.13 ) | – – on the 16 basis probabilities | 12 | – | 145 | 157 |
| B32: boundary controls (App. B.10 ), nonlinear projection + magnitude readout | on | 12 | 48 † | 33 | 93 |
| Method | Params | Avg. test | vs. SQUARE [avg. taskwise interval] |
| LoRA | 417 | 0.5897 | [ ] |
| MLP | 1,633 | 0.6817 | [ ] |
| SQUARE w/o PQC | 145 | 0.6412 | [ ] |
| SQUARE, frozen circuit | 145 | 0.6155 | [ ] |
| SQUARE | 157 | 0.7565 | – |
| Givens-12 + squares + head | 157 | 0.7271 | [ ] |
| Pre-entangler rotations | Entangler | Angles | Total | Avg. test |
| none | 8 | 153 | 0.7891 | |
| linear CNOT | 8 | 153 | 0.7879 | |
| ring CNOT | 8 | 153 | 0.7568 | |
| none | 12 | 157 | 0.7889 | |
| linear CNOT | 12 | 157 | 0.7875 | |
| ring CNOT (default) | 12 | 157 | 0.8462 |
| Experiment | Data Source | Train | Validation | Test / Eval | Seeds | Evaluation Unit |
| 2D decision-boundary analysis (nominal) | Controlled 2D bottleneck features | 1,000 | 400 | – | 3 | Binary example |
| GLUE-derived interaction classification | GLUE inputs with controlled nonlinear labels | 7,600 | 2,250 | – | 3 | Sentence / sentence-pair example |
| Alpaca candidate reranking | Alpaca instruction-following candidates | 960 | 120 | 120 | 3 | Query–candidate pair |
| Disjoint held-out control suite | Eight GLUE inputs with controlled nonlinear labels | Disjoint task-specific subsets | 5 | Sentence / sentence-pair example | ||
| Item | Value |
| Train examples | 1,000 |
| Validation examples | 400 |
| Input dimension | 2 |
| Label construction | Controlled nonlinear rule |
| Evaluation unit | Binary example |
| Metrics | Accuracy, F1, AUC |
| Item | Value |
| Train queries | 960 |
| Validation queries | 120 |
| Test queries | 120 |
| Candidates per query | 8 |
| Evaluation unit | Query–candidate pair |
| Selection rule |
| Method | Hyperparameter | Values |
| BitFit | Batch Size | |
| Optimizer | AdamW | |
| Scheduler | Linear Scheduler | |
| Learning Rate | ||
| Trainable Parameters | Bias terms only | |
| LoRA | Batch Size |