Wi-Fi sensing promises to turn the everyday wireless signals that already surround us into ubiquitous sensors for human sensing. However, a fundamental obstacle is that CSI is acquired under diverse device-specific configurations, including different subcarrier counts, bandwidths, and carrier bands. Consequently, the resulting CSI tensors vary in both spectral resolution and tensor shape, making heterogeneous modeling challenging. Standard architectures struggle with such heterogeneity, forcing lossy pre-processing which compromises the underlying signal. To bridge this gap, we present UniCSI, a unified foundation architecture that directly operates on heterogeneous CSI while preserving the integrity of the native waveform. UniCSI hinges on two core innovations: (1) a physics-informed RF tokenizer that encodes each frequency channel based on its fractional position within the physical spectrum rather than rigid array indices. It preserves intrinsic spectral coherence and enables seamless, frequency resolution-agnostic processing across arbitrary sensing configurations. (2) a spectral aggregator that distills variable-length channel sequences into a fixed-size spectral signature, effectively decoupling the feature dimensionality from the physical subcarrier spacing. Extensive evaluations on a large-scale corpus of 25 heterogeneous public datasets, spanning 14 to 2048 subcarriers, 20 to 160 MHz bandwidth, and the 2.4 and 5 GHz bands, demonstrate that native heterogeneous ingestion substantially improves cross-domain transfer under both supervised and self-supervised training schemes, particularly in regimes where fixed-grid architectures fail to generalize.
Figures & tables
Figure 1: Overview of UniCSI. (Top) Per-link encoder: a temporal CNN and time-patch projection tokenize per-subcarrier CSI; learned queries with band-relative positional encoding cross-attend into the subcarrier tokens, compressing the variable-length spectrum into a fixed Q×P×d latent for a Transformer encoder. (Bottom) Training regimes: JEPA pretraining, supervised multi-task training, and downstream probing or fine-tuning.
Dataset
Task ( n )
Rx
Tx
Sub.
Band (GHz)
BW (MHz)
Wi-Fi Sensing Contest ( Han 2024 )
Presence ( n=3 )
2
2
248, 250
5
80, 160
Person-in-Wi-Fi 3D ( Yan et al. 2024 )
3D Pose ( n=4 )
3×3
1
30
5
20
XRFV2 ( Lan et al. 2025 )
HAR ( n=30 )
3×3
1
30
5
20
Wi-MIR ( Islam et al. 2024 )
HAR ( n=17 )
3
3
30
5
20
Behavior Auth. ( Shi et al. 2017 )
User ID ( n=12 )
3
1
30
5
20
NTU-Fi Human ID ( Yang et al. 2022 )
User ID ( n=14 )
3
1
114
5
40
Table 1: CSI datasets for pretraining and evaluation. Shaded rows form the self-supervised pretraining corpus; unshaded rows are held out for downstream evaluation only. n is the number of class labels, Rx/Tx the antenna counts, Sub. the subcarrier count, Band the carrier band and BW the channel bandwidth; D×A denotes D receiver devices of A antennas each. Footnote marks are explained in the text.
Multi-task
JEPA
Probe / k -NN
Fine-tune
Encoder
scratch
EMA
frozen
fine-tuned
Head
linear
predictor
BN + lin.
BN + lin.
Objective
CE
smooth- L1
CE / cosine
CE
LR
3×10−4
1×10−4
1×10−3
3×10−4
WD
5×10−2
5×10−2
1×10−5
5×10−2
Budget
50 epochs
50 k steps
200 epochs
50 epochs
Table 2: Training hyperparameters. All runs use AdamW, cosine decay to 0 , gradient clip 1.0 , and bf16.
Figure 2: Fused F1 of a single encoder trained jointly across eight heterogeneous CSI datasets. Bars denote the mean and whiskers the standard deviation over three seeds. All datasets except NTU-Fi HAR (in-distribution control) use a cross-environment or cross-subject protocol.
Figure 3: Linear probe and k -NN evaluation of Padded ViT and UniCSI on eight representative dataset splits. Bars report fused macro-F1; Avg is the mean over the selected splits.
Dataset
Split
Scratch (Ours)
Full FT (PaddedViT)
Full FT (Ours)
WiMANS (5 GHz)
Cross-Env
55.69%
36.96%
55.83%
XRF55
Cross-Env
39.53%
15.5%
39.78%
SHARP
Cross-Env
58.50%
75.71%
71.06%
Exposing-CSI
Cross-Env
47.56%
7.6%
45.16%
FallDar
Random
96.78%
98.61%
98.95%
CSI-Bench (2.4 GHz)
Random
80.19%
91.16%
90.67%
Table 3: Comparison between supervised training from scratch with UniCSI, full fine-tuning (FT) from the JEPA-pretrained Padded ViT, and full FT from the JEPA-pretrained UniCSI on unseen downstream datasets (Metric: Macro F1 score).
Figure 4: Few-shot downstream fine-tuning from pretrained versus random initialization, averaged over eight dataset–split pairs.
WiFi sensing based on Channel State Information (CSI) promises ubiquitous, device-free perception, yet current research remains trapped in a Tower of Babel - fragmented into isolated silos where models are tailored to specific hardware dialects, fixed environments, and narrow tasks. The primary bottleneck is the Heterogeneity Gap: the disparity in signal dimensions, sampling rates, and semantic labels that prevents cross-system understanding. To bridge this gap, we propose a foundation-model framework that treats CSI not merely as raw signals but as a structured language with a learnable universal grammar. We first curate and standardize a large collection of heterogeneous real-world CSI datasets, establishing a unified infrastructure that allows incompatible signal formats to be treated as a single corpus. Second, we introduce a modular architecture that acts as a universal translator where lightweight dataset-specific adapters tokenize diverse signal inputs into a shared latent vocabulary, while a shared self-supervised Transformer backbone learns the temporal syntax of human motion and environmental dynamics. This design decouples sensing semantics from hardware syntax. Extensive evaluations show that by mastering this universal language, our approach consistently outperforms task-specific baselines and exhibits strong generalization capability in new environments, achieving superior efficiency in few-shot scenarios. By effectively absorbing heterogeneity, the framework offers a path toward robust, general-purpose wireless sensing, mirroring the linguistic generalization observed in Large Language Models. The code implementation is available at: https://github.com/cjychenjiayi/WiLLM.
Jiayi Chen, Weiting Ou, Guangxu Zhu
Shenzhen Research Institute of Big Data · The Chinese University of Hong Kong, Shenzhen · Shenzhen Loop Area Institute (SLAI)
Existing Wi-Fi sensing systems rely on injecting high-rate probing packets to extract channel state information (CSI), leading to communication degradation and limited deployment flexibility. Although Integrated Sensing and Communication (ISAC) is a promising direction, existing solutions still rely on auxiliary packet injection because they exploit only uniform CSI from a single frame type, discarding approximately 70% of naturally available packets. We present FuseFi, a novel Wi-Fi-based ISAC framework that directly exploits irregularly sampled CSI from diverse communication packets across multiple frequency bands, eliminating intrusive packet injection and introducing no sensing-specific communication overhead. FuseFi integrates a CSI sanitization pipeline to harmonize heterogeneous packets and remove burst-induced redundancy, together with a time-aware attention model that learns directly from non-uniform CSI sequences without resampling. We further introduce CommCSI-HAR, a new dataset with irregularly sampled CSI from real-world dual-band communication traffic. Extensive evaluations on this dataset and, separately, on four public sensing tasks across three benchmark datasets show that FuseFi achieves state-of-the-art accuracy with a compact model size, while fully preserving communication throughput.
Gaofeng Dong, Kang Yang, Mani Srivastava
Department of Electrical and Computer Engineering, University of California, Los Angeles, CA, USA
Channel state information (CSI) provides a widely available sensing modality for human and environment perception, but existing CSI sensing models usually rely on task-specific supervised training and require substantial labeled data for each task, device, user, or environment. This limits their scalability in practical deployments where unlabeled CSI is abundant but labeled data is costly to collect. In this paper, we present CSI-JEPA, a self-supervised predictive representation learning framework for label-efficient, multi-task Wi-Fi sensing. CSI-JEPA learns reusable temporal-spectral representations from unlabeled CSI samples by predicting latent features of masked channel regions from visible context. To better match the physical structure of CSI, CSI-JEPA tokenizes channel-response amplitude windows along the time and subcarrier dimensions. It then introduces a channel variation-aware masking strategy that samples predictive targets from regions with stronger local temporal and subcarrier-domain variations. After pretraining, the encoder is frozen and used as a backbone, with lightweight task-specific adapters added for downstream sensing tasks. We evaluate CSI-JEPA on seven real-world Wi-Fi sensing tasks spanning diverse objectives and deployment settings. The results show that CSI-JEPA improves downstream sensing performance over competitive baselines, achieving up to 10.64 percentage points mean accuracy gain over state-of-the-art supervised Transformer and matched-budget label savings of up to 98.0%.