Qubit-centric Transformer for Surface Code Decoding
Authors: Seong-Joon Park, Hee-Youl Kwak, Yongjune Kim
Organizations: Samsung Electronics Company, Ltd., Suwon 16677, South Korea · Institute of Artificial Intelligence, Pohang University of Science and Technology (POSTECH), Pohang 37673, South Korea · Department of Electrical, Electronic and Computer Engineering, University of Ulsan, Ulsan 44610, South Korea · Department of Electrical Engineering, Pohang University of Science and Technology (POSTECH), Pohang 37673, South Korea
For reliable large-scale quantum computation, quantum error correction (QEC) is essential to protect logical information distributed across multiple physical qubits. Taking advantage of recent advances in deep learning, neural network-based decoders have emerged as a promising approach to improve the reliability of QEC. We propose the qubit-centric transformer (QCT), a novel and universal QEC decoder based on a transformer architecture with a qubit-centric attention mechanism. Our decoder transforms input syndromes from the stabilizer domain into qubit-centric tokens via a specialized embedding strategy. These qubit-centric tokens are processed through attention layers to effectively identify the underlying logical error. Furthermore, we introduce a graph-based masking method that incorporates the topological structure of quantum codes, enforcing attention toward relevant qubit interactions. Across various code distances for surface codes, QCT achieves state-of-the-art decoding performance, significantly outperforming existing neural decoders and the belief propagation (BP) with ordered statistics decoding (OSD) baseline. Notably, QCT achieves a high threshold of 18.1% under depolarizing noise, which closely approaches the theoretical bound of 18.9% and surpasses both the BP+OSD and the minimum-weight perfect matching (MWPM) thresholds. This qubit-centric approach provides a scalable and robust framework for surface code decoding, advancing the path toward fault-tolerant quantum computing.
Figures & tables
Figure 1: Illustration of the topological structure of the surface code with distance d=3 (left) and the syndrome extraction circuits (right).
Figure 2: Illustration of the high-level decoding methodology treating decoding as a classification problem. An arbitrary physical error E can be decomposed into E=S⋅T⋅L .
Figure 3: The proposed QCT architecture. QCT processes the input syndrome vector s through the embedding layer and merging layer, followed by N transformer blocks, to finally predict the logical operator.
Figure 4: Illustration of the embedding process for physical qubit q5 in a d=3 surface code.
Figure 5: Illustration of the merging layer. The input consists of n decoupled Z-tokens {ϕiZ}i=1n (yellow) and n decoupled X-tokens {ϕiX}i=1n (blue) separately. This process fuses the decoupled embeddings into unified qubit tokens x(0) .
Figure 6: Illustration of the structure-aware masks M for surface codes used in QCT. The left figure shows the mask for d=3 surface code, and the right figure shows the mask for d=5 surface code.
Figure 7: Performance under the code-capacity noise model. Comparison of the proposed QCT decoder against MWPM, FFNN, and CNN decoders for code distances (a) d=5 , (b) d=7 , and (c) d=9 . The MWPM, FFNN, and CNN results are taken from [ 23 ] .
Figure 8: Performance under the code-capacity noise model. Comparison of the proposed QCT decoder against the BP+OSD baseline for code distances d∈{5,7,9,11,13} . At least 1000 logical errors are collected for d∈{5,7,9} and at least 500 for d∈{11,13} . Error bars indicate 95% Wilson confidence intervals. The values at p=0.01 and p=0.02 are obtained by extrapolation and are shown without markers.
Figure 9: Error suppression factor Λ under the code-capacity noise model, obtained from the LERs at p=0.05 for QCT, BP+OSD, and MWPM. At least 1000 logical errors are collected for d∈{5,7,9} and at least 500 for d∈{11,13} . Error bars on the LER data indicate 95% Wilson confidence intervals.
Figure 10: Performance under the circuit-level noise model in terms of the LER normalized per round. At least 1000 logical errors are collected for each simulated data point, and error bars indicate 95% Wilson confidence intervals. The results at p=0.001 are obtained by extrapolation and are shown without markers.
Figure 11: Error suppression factor Λ5/7 under the circuit-level noise model for QCT, standard MWPM, BP+OSD, and correlated MWPM. At least 1000 logical errors are collected for each simulated LER used to compute Λ5/7 ; the values at p=0.001 are based on extrapolated LERs and are shown without markers.
Figure 12: Decoding performance of QCT under various configurations: (a) impact of the number of transformer blocks ( N ), (b) effect of structure-aware masking, and (c) effect of the merging layer to decoding accuracy. At least 500 logical errors are collected for each simulated data point, and error bars indicate 95% Wilson confidence intervals.
d
3
5
7
9
11
13
Code-capacity
–
1.61 M
1.61 M
1.62 M
1.63 M
1.64 M
Circuit-level
2.00 M
2.02 M
2.05 M
–
–
–
Table 1: Number of trainable parameters for the evaluated surface-code distances.
d
3
5
7
MWPM [ 10 ]
8.28%
10.36%
11.94%
FFNN [ 41 ]
9.77%
11.35%
12.49%
CNN [ 23 ]
9.80%
12.15%
13.26%
QCT (Proposed)
9.80%
13.27%
14.31%
Table 2: Comparison of the pseudo-thresholds for the proposed QCT against benchmark decoders (MWPM, FFNN, and CNN) on the surface code for distances d=3,5, and 7 under the code-capacity noise model. The results of MWPM, FFNN, CNN are from [ 23 ] .
Figure 13: Threshold performance for surface codes with d=5,7,9,11,13 . (a) QCT. (b) BP+OSD. At least 1000 logical errors are collected for d∈{5,7,9} and at least 500 for d∈{11,13} . Error bars indicate 95% Wilson confidence intervals.
Real-time decoding is a major bottleneck in scaling quantum error correction (QEC) from noisy intermediate-scale quantum (NISQ) devices to fault-tolerant quantum computing. We present an adaptive confidence-gated decoding framework for the rotated surface code that treats decoding as a two-stage inference problem. A lightweight feed-forward neural network performs fast-path decoding for the majority of syndrome measurements, while only low-confidence predictions are escalated to a minimum-weight perfect matching (MWPM) refinement stage. We benchmark the framework on rotated surface codes with distances d∈{3,5,7,9,11} under circuit-level depolarising noise using the Stim stabiliser simulator. The evaluation characterises logical accuracy, confidence-controlled accuracy-latency trade-offs, decoding throughput, per-shot latency, and decoding-graph resource scaling. Routing only 3.3%-6.2% of syndromes to the refinement stage improves logical accuracy from 99.21% for the neural-only baseline to 99.81% at a confidence threshold of 0.95 while incurring only a bounded increase in average decoding cost. Neural-decoder throughput saturates near 4.6×105 samples s−1 at batch size 512 on commodity CPU hardware, indicating that the neural fast path is not the dominant throughput bottleneck beyond code distance d=7. We release the complete benchmarking pipeline, trained models, raw benchmark data, and source code, and explicitly distinguish the experimentally validated contributions from the broader hardware-aware QEC co-design roadmap, including hardware-constrained code discovery, GPU-accelerated inference, and multi-noise optimisation, which remain directions for future work.
Sumit Chongder
Department of Physics, Quantum Information and Computation, Indian Institute of Technology Jodhpur, Jodhpur, Rajasthan 342037, India
Foundation decoders, a class of high-capacity neural decoders, are leading candidates for fault-tolerant quantum computing, with accurate and efficient decoding at large code distances. However, their construction often faces a steep scaling barrier, as larger code distances rapidly amplify the cost of syndrome generation and neural optimization. To address this bottleneck, here we devise neural transfer unification (NTU), a unified framework for efficient foundation decoders. A central feature of NTU is its ability to align decoding tasks across code distances via algebraic structures shared by scalable code families, which enables knowledge learned on smaller codes to accelerate large-scale decoder training. We instantiate NTU as NTU-Transformer, a transformer-based neural decoder tailored for planar surface codes and bivariate bicycle codes. For planar surface codes under circuit-level noise, NTU-Transformer outperforms correlation-aware matching on the [[361,1,19]] code and further scales to the [[625,1,25]] code, where it exceeds standard matching through transfer adaptation. For the bivariate bicycle code with [[72,12,6]], it surpasses Relay-BP in the low-physical-error regime. These results establish our proposal as a scalable route to amortized cross-distance training of foundation decoders for fault-tolerant quantum processors.
Ge Yan, Shanchuan Li, Shiyi Xiao +4
College of Computing and Data Science, Nanyang Technological University, Singapore · Department of Electrical Engineering and Computer Science, Tokyo University of Agriculture and Technology, Koganei, Tokyo, Japan · School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China +2
Quantum error correction (QEC) is essential for enabling quantum advantages, with decoding as a central algorithmic primitive. Owing to its importance and intrinsic difficulty, substantial effort has been made to QEC decoder design, among which neural decoders have recently emerged as a promising data-driven paradigm. Despite this progress, practical deployment remains hindered by a fundamental accuracy-latency tradeoff, often on the microsecond timescale. To address this challenge, here we revisit neural decoders for surface-code decoding under explicit accuracy-latency constraints, considering code distances up to d=9 (161 physical qubits). We unify and redesign representative neural decoders into five architectural paradigms and develop an end-to-end compression pipeline to evaluate their deployability and performance on FPGA hardware. Through systematic experiments, we reveal several previously underexplored insights: (i) near-term decoding performance is driven more by data scale than architectural complexity; (ii) appropriate inductive bias is essential for achieving high decoding accuracy; and (iii) INT4 quantization is a prerequisite for meeting microsecond-scale latency requirements on FPGAs. Together, these findings provide concrete guidance toward scalable and real-time neural QEC decoding.
Ge Yan, Shanchuan Li, Yuxuan Du
College of Computing and Data Science, Nanyang Technological University, Singapore 639798, Singapore · Department of Electrical Engineering and Computer Science, Tokyo University of Agriculture & Technology, Koganei, Tokyo, 184-8588, Japan · School of Physical and Mathematical Sciences, Nanyang Technological University, Singapore 639798, Singapore