Decoding signals over unknown channels with minimal pilot overhead is a critical challenge in next-generation communications. Existing deep learning approaches typically rely on generic encoders that struggle to model long-range temporal dependencies or efficiently capture the channel's physical properties from scarce data. We argue that standard architectures suffer from agnostic estimation gaps, as they must implicitly learn the constellation geometry that is already known. We introduce the Constellation-Aware Transformer (CAT), a novel architecture that explicitly injects geometric inductive biases into the equalization process. CAT is composed of a stack of custom TransFIRmer blocks, which use an "early interaction" paradigm to co-process received signals and ideal constellation symbols. Each block features a split Feed-Forward Network that applies a Finite Impulse Response (FIR)-inspired filter for deconvolution and a parallel MLP for geometric refinement. We show that this design is structurally aligned with the optimal linear (MIMO Wiener) receiver: its attention can implement a matched-filter bank, and its bidirectional FIR branch provides the non-causal filtering that block MMSE equalization requires. In the semi-supervised setting, CAT needs fewer pilots than VAE and standard Transformer baselines: on two of our three ISI channels, it reaches a lower SER with 64 pilots than they do with 128.
Figures & tables
Modulation
K
Attn. Overhead
SER Gain vs. Vanilla Trans.
16-QAM
16
∼ 13%
3.8 dB
64-QAM
64
∼ 56%
3.0 dB
256-QAM
256
∼ 300%
2.2 dB
Table 1 : Scalability across constellation orders ( h(1) channel, N=256 , SNR 16–24 dB, gains at target SER 10−2 ).
Channel
Pilots
VAE-CNN
Trans.
CAT (Ours)
Optimal
h(1) (L=5)
16
0.2900
0.3392
0.3580
0.0121
32
0.1251
0.1192
0.0842
64
0.0523
0.0610
0.0198
128
0.0494
0.0290
0.0156
h(2) (L=4)
16
0.3447
0.3563
0.3330
0.0101
32
0.1843
0.1593
0.1372
Table 2 : SER on channels with memory ( Ex/N0=17 dB, payload 256). CAT vs. VAE-CNN and Vanilla Transformer.
Method
14 dB
17 dB
20 dB
24 dB
26 dB
VAE-CNN
0.1842
0.0523
0.0215
0.0068
0.0031
Vanilla Trans.
0.1763
0.0610
0.0248
0.0091
0.0042
CAT (Ours)
0.1205
0.0198
0.0072
0.0019
0.0008
Table 3 : SER vs. SNR on h(1) ( Np=64 , payload 256). The 17 dB column repeats Table 2 .
Model Variant
Memoryless
ISI ( h(1) )
Our Full Method
CAT (built with TransFIRmer blocks)
0.0599 ± 0.0005
0.0198
Architecture & Prior Ablations
CAT without Inverse FIR Filter
0.0608 ± 0.0005
0.0287
CAT with MLP-FFN (No FIR)
0.0615 ± 0.0006
0.0351
No Prior + Bidirectional FIR
–
0.0438
Table 4 : Ablation study. Memoryless: Np=16 , SNR =22 dB, 64 payload symbols (as in Figure 2(b) ), SER ± 95% CI. ISI: h(1) , Ex/N0=17 dB, Np=64 , payload 256. 4 4 4 “Vanilla Trans” = self-attention on signals only; “Self-Only” = signals and constellations attend within their own group; “w/o Inverse FIR” = forward conv only; “No Prior + Bidirectional FIR” = signal-only attention with the TransFIRmer FFN; “Pos. Emb. on Constellation” = positional embeddings also added to the constellation tokens.
Method
BLER Gain
ECE
Calibration Improvement
VAE-CNN
—
0.085
—
Vanilla Transformer
0 dB (reference)
0.089
—
CAT (Ours)
+2.5 dB
0.012
7.4×
Table 5 : Coded BLER with 3GPP 5G NR LDPC (Rate 1/2, CW=1024) on the h(1) channel. Gains at target BLER =10−2 .
Profile
Delay spread
Vanilla Trans.
VAE-CNN
CAT (Ours)
Gain
TDL-A (NLOS)
30 ns
0.0385
0.0421
0.0195
1.97×
TDL-D (LOS)
30 ns
0.0198
0.0231
0.0112
1.77×
Table 6 : SER on 3GPP TDL channels (16-QAM, N=128 , 500 realizations per profile). Gain is the SER of the stronger baseline divided by that of CAT.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
VAE-CNN
Vanilla Trans.
CAT (Ours)
Parameter Count
∼150 k
∼9 k
∼10 k
Self-Attn Ops
0
O(N2d)
O((N+K)2d)
Pilot Efficiency
Low
Medium
High
Appendix
Table 7 : Parameter and complexity comparison (approximate for N=128 , d=10 ).
Channel
SNR
CAT
CAT w/ MLP-FFN
FIR gain
h(1) (L=5)
14 dB
0.1205
0.1284
7%
17 dB
0.0198
0.0351
77%
24 dB
0.0019
0.0025
32%
h(2) (L=4)
14 dB
0.0921
0.1195
30%
17 dB
0.0340
0.0489
44%
24 dB
0.0024
0.0041
71%
Appendix
Table 8 : CAT vs. CAT with MLP-FFN on ISI channels ( Np=64 , payload 256). The CAT entries for h(1) are those of Table 3 . FIR gain is the relative SER increase when the FIR-FFN is replaced by an MLP.
Steps
Method
SER
SER increase
5,000
Cold-start
1.56×10−2
—
200
Cold-start
1.63×10−2
+4%
50
Cold-start
2.30×10−2
+47%
50
CAVIA-init
1.59×10−2
+2%
10
CAVIA-init
1.82×10−2
+17%
Appendix
Table 9 : Convergence speed under different adaptation strategies ( h(1) , Ex/N0=17 dB, Np=128 pilots, payload 256). The 5,000-step cold-start run is the Np=128 result of Table 2 ; the last column is the SER increase relative to it.
Setting
VAE-CNN
Vanilla Trans.
CAT
Overhead
N =128, 16-QAM
11.2 μ s
12.5 μ s
16.0 μ s
+28%
N =256, 16-QAM
18.1 μ s
21.0 μ s
24.5 μ s
+17%
N =128, 64-QAM
11.2 μ s
12.5 μ s
24.8 μ s
+98%
Appendix
Table 10 : Wall-clock time of one training step (forward pass, backward pass and optimizer update; single RTX 4090, batch size 1).