Organizations: University of Science and Technology Beijing, China · Department of ECE, The Hong Kong University of Science and Technology, Hong Kong · University of Exeter · Sony China Research Laboratory, China
With the growing demand for privacy-preserving and occlusion-resilient human pose recognition (HPR), 5G channel state information (CSI) offers a promising contactless sensing modality by integrating communication and sensing capabilities. However, collecting large-scale synchronized CSI-pose pairs remains costly in practical 5G systems. To address this limitation, we propose StructFlow-HPR, a structured pose-conditioned flow matching framework for generative CSI augmentation. StructFlow-HPR learns a continuous latent transport process from Gaussian noise to real CSI representations under pose guidance, while preserving the receiver-frequency topology of CSI through a reconstruction-preserving autoencoder. A pose-conditioned Transformer is further designed to model the latent velocity field and generate pose-aligned CSI samples via ordinary differential equation sampling. Experiments on real-world 5G sensing data show that StructFlow-HPR can produce realistic CSI-pose pairs and improve downstream HPR performance under limited-data conditions.
Figures & tables
Fig. 1: Overall framework of StructFlow-HPR for pose-conditioned CSI generation and downstream HPR augmentation.
State
PCK5
PCK10
PCK20
PCK30
A1
A2
A3
Δ
A1
A2
A3
Δ
A1
A2
A3
Δ
A1
A2
A3
Δ
Squat
67.84
76.93
77.79
+0.86
85.67
94.42
95.45
+1.03
91.46
99.77
99.84
+0.07
91.83
100.00
100.00
+0.00
Move
76.61
85.78
86.35
+0.57
89.58
98.39
99.08
+0.69
91.73
99.98
99.95
-0.03
92.16
100.00
100.00
+0.00
Rise Hand1
90.39
99.22
99.43
+0.21
91.27
99.91
99.91
+0.00
92.14
100.00
100.00
+0.00
91.76
100.00
100.00
+0.00
Rise Hand2
67.18
76.31
77.22
+0.91
84.53
93.29
94.14
+0.85
90.62
99.31
99.31
+0.00
91.49
99.84
99.68
-0.16
Press Leg1
43.82
52.99
53.63
+0.64
79.06
88.03
88.45
+0.42
88.74
97.31
97.13
-0.18
88.93
97.70
97.82
+0.11
TABLE I: Performance comparison between MetaFi, MfDfHPR, and StructFlow-HPR under various PCK thresholds (Motion States).
Fig. 2: t-SNE visualization of structured CSI latent distributions for real and StructFlow-generated samples under representative motion states.
Fig. 3: The human pose coordinates are generated by the visual model and our proposed StructFlow-HPR model, respectively.
Accurate channel state information (CSI) prediction is essential for proactive beamforming and resource management in 5G massive MIMO systems, yet the deployment of high-accuracy transformer-based predictors on base-station hardware remains challenging because the most capable models carry upwards of 30,M parameters. This paper introduces Lightweight PCGAE-Net, which addresses the efficiency problem not by post-hoc compression but by correcting two architectural flaws in the current state of the art. The first is a sequential attention ordering bias: in CS3T-UNet, group-wise temporal attention (GTA) always operates on features that have already been transformed by cross-shaped spatial attention (CSA), distorting what temporal information GTA can capture. We remove this dependency by routing both attention modules to the same layer-normalized input and combining their independent outputs through a learned per-channel sigmoid CrossGate. The second flaw is an uncompressed bottleneck: applying full self-attention at the deepest encoder stage, where channel depth reaches 4C, is quadratically expensive and carries redundant features. A Bottleneck AutoEncoder (BAE) with 1×1 convolutions halves this depth and uses an auxiliary reconstruction loss to prevent information collapse. Wrapping these components inside a shallower encoder-decoder with frequency-domain dimensionality reduction (Nf=32, C=48) produces a model with just 8.54,M parameters -- 58% fewer than the CS3T-UNet baseline -- that outperforms it by up to 3.26,dB at 5,km/h and 6.0,dB at 9,km/h in single-step prediction on QuaDriGa dataset.
Uma Kishore Godavarti, K. Giridhar, Vanani Prince Dharmendrabhai +2
Network Modem Team, Samsung R&D India Bangalore · Department of Electrical Engineering, Indian Institute of Technology Madras, Chennai, India
WiFi Channel State Information (CSI) enables privacy-preserving human pose sensing in camera-denied environments, but existing WiFi-based pose estimators often fail under environment shifts and rely on costly camera-based annotation pipelines that limit scale. We propose WiFi-JEPA, a self-supervised framework that learns CSI-native representations by predicting masked latent embeddings instead of reconstructing raw CSI signals that may contain hardware-specific artifacts. WiFi-JEPA makes three contributions: (i) CSI-specific tokenization and link masking tailored to the CSI tensor over channel, time, and link (C,T,L); masking entire Tx-Rx antenna links forces the model to predict one spatial link view from others, capturing cross-link correlations informative of 3D spatial structure. (ii) A ray-tracing CSI simulation pipeline that generates diverse unlabeled CSI from randomized geometric primitives, providing scalable pre-training data without pose annotations. (iii) State-of-the-art results on Person-in-WiFi-3D: WiFi-JEPA outperforms prior WiFi-CSI baselines on both single- and multi-person 3D pose estimation under the same evaluation protocol. We also show that simulated CSI provides complementary pre-training signal to real CSI, and that four vision-native SSL objectives degrade performance below training from scratch, whereas WiFi-JEPA consistently improves downstream pose estimation.
Doeon Kim, Jungyoon Lee, Seongsin Kim +1
Department of Intelligent Semiconductors, Soongsil University, Republic of Korea · Department of AI Convergence Security, Soongsil University, Republic of Korea · School of AI Software, Soongsil University, Republic of Korea
Accurate yet low-latency channel state information (CSI) acquisition is essential for multiple-input multiple-output (MIMO) communication systems. While advanced deep generative models, such as score-based and diffusion models, enable high-fidelity CSI reconstruction from limited pilot observations, they often suffer from high inference latency. To achieve accurate CSI estimation under stringent latency constraints, this paper proposes a null-space flow matching (FM) framework that decomposes pilot-limited MIMO channel estimation into a range-space reconstruction problem and a null-space generation problem. Specifically, the range-space component of the channel is directly recovered from noisy pilot observations, while only the ambiguous null-space component is iteratively refined using an FM-based generative prior. To further improve the robustness of the proposed framework, we introduce a power-law time schedule to better allocate the limited number of refinement steps, along with a noise-aware adaptive correction strategy to suppress channel noise on the refinement trajectory. Experimental results demonstrate that our method achieves a competitive normalized mean square error (NMSE) even under a strict latency budget of around 3 ms, while delivering superior estimation accuracy and faster inference than both model-based and generative baselines.
Junjie Zhao, Guangming Liang, Dongzhu Liu +1
College of Electronic Information and Optical Engineering, Nankai University · School of Computing Science, University of Glasgow · School of Natural and Computing Science, University of Aberdeen