cs.LGSep 27, 2026

Geometric Inductive Biases for Semi-Supervised Equalization: The Constellation-Aware Transformer

Authors: Avi Caciularu

Organizations: Google Research

Abstract

Decoding signals over unknown channels with minimal pilot overhead is a critical challenge in next-generation communications. Existing deep learning approaches typically rely on generic encoders that struggle to model long-range temporal dependencies or efficiently capture the channel's physical properties from scarce data. We argue that standard architectures suffer from agnostic estimation gaps, as they must implicitly learn the constellation geometry that is already known. We introduce the Constellation-Aware Transformer (CAT), a novel architecture that explicitly injects geometric inductive biases into the equalization process. CAT is composed of a stack of custom TransFIRmer blocks, which use an "early interaction" paradigm to co-process received signals and ideal constellation symbols. Each block features a split Feed-Forward Network that applies a Finite Impulse Response (FIR)-inspired filter for deconvolution and a parallel MLP for geometric refinement. We show that this design is structurally aligned with the optimal linear (MIMO Wiener) receiver: its attention can implement a matched-filter bank, and its bidirectional FIR branch provides the non-causal filtering that block MMSE equalization requires. In the semi-supervised setting, CAT needs fewer pilots than VAE and standard Transformer baselines: on two of our three ISI channels, it reaches a lower SER with 64 pilots than they do with 128.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 6, 2026eess.SP

Learning the Channel Gain from Anywhere to Anywhere via Cross-environment Transformer Estimators

Channel-gain maps provide the channel gain between any two locations in a geographical region. They find numerous applications, from resource allocation and interference control to path planning for autonomous vehicles. Channel-gain map estimation (CGME) is considerably more challenging than conventional radio map estimation (RME) because channel-gain maps are functions over a 6-dimensional input space. This calls for specialized methods, which currently rely on the (inaccurate) radio tomographic model or require a prohibitively large number of measurements since they do not exploit any spatial structure. This paper overcomes this issue by leveraging spatial patterns that channel-gain maps exhibit across environments, as dictated by the laws of physics and typical environmental characteristics (e.g. building materials and layouts). Adopting a metalearning perspective, a transformer-based estimator is proposed to implicitly learn this common structure from measurements collected in multiple environments. This enables CGME in new environments from significantly fewer measurements (five times less in our experiments). To maximize learning efficiency, the transformer is composed with a feature map that enforces the invariances of CGME, such as those following from reciprocity. Numerical experiments corroborate the merits of the proposed estimator relative to existing methods.
May 19, 2026eess.SP

PilotWiMAE: Pilot-Native Representation Learning for Wireless Channels

Channel foundation models assume access to fully observed channels, an assumption that fails in deployment. We introduce PilotWiMAE, a self-supervised framework whose encoder ingests noisy pilot observations directly and whose attention factorizes along the axis separating temporal from joint space-frequency processing, an inductive bias inspired by the physics of the problem. Pilot input shrinks the observation space by up to two orders of magnitude and also removes the unrealistic assumption of full-CSI availability while incurring lower latency. The factorized design generates robust representations by exploiting the separable channel structure and allows a pretraining mask ratio of 99%99\%. We pair patch-normalized reconstruction, which captures small-scale fading structure, with an auxiliary scale loss that recovers the large-scale fading features, and use an AWGN curriculum to match pilot noise at pretraining and deployment. Pretrained solely on 3.53.5,GHz and evaluated at 2828,GHz across in-distribution and out-of-distribution settings, PilotWiMAE's cross-frequency beam selection and channel characterization beat supervised baselines despite operating on a smaller observation space. To weaken the coupling between decoder capacity and representation quality, we further propose a decoder-centric pretraining stage following the encoder-decoder joint pretraining, which allows PilotWiMAE to demonstrate competitive channel estimation without sacrificing representation quality. To foster further work in this direction, we release the PilotWiMAE pretrained weights and training pipeline, together with CSIGen, our Sionna-based ray-tracing channel-generation tool, and the channel datasets used in this work.
Jun 24, 2026cs.NI

Lightweight PCGAE-Net: Parallel CrossGate Attention and Bottleneck AutoEncoder for Efficient 5G Channel Prediction

Accurate channel state information (CSI) prediction is essential for proactive beamforming and resource management in 5G massive MIMO systems, yet the deployment of high-accuracy transformer-based predictors on base-station hardware remains challenging because the most capable models carry upwards of 30,M parameters. This paper introduces Lightweight PCGAE-Net, which addresses the efficiency problem not by post-hoc compression but by correcting two architectural flaws in the current state of the art. The first is a sequential attention ordering bias: in CS3T-UNet, group-wise temporal attention (GTA) always operates on features that have already been transformed by cross-shaped spatial attention (CSA), distorting what temporal information GTA can capture. We remove this dependency by routing both attention modules to the same layer-normalized input and combining their independent outputs through a learned per-channel sigmoid CrossGate. The second flaw is an uncompressed bottleneck: applying full self-attention at the deepest encoder stage, where channel depth reaches 4C4C, is quadratically expensive and carries redundant features. A Bottleneck AutoEncoder (BAE) with 1×11\times1 convolutions halves this depth and uses an auxiliary reconstruction loss to prevent information collapse. Wrapping these components inside a shallower encoder-decoder with frequency-domain dimensionality reduction (Nf ⁣= ⁣32N_f\!=\!32, C ⁣= ⁣48C\!=\!48) produces a model with just 8.54,M parameters -- 58% fewer than the CS3T-UNet baseline -- that outperforms it by up to 3.26,dB at 5,km/h and 6.0,dB at 9,km/h in single-step prediction on QuaDriGa dataset.