TRACE: Learning to Self-Calibrate Wireless Digital Twins from ISAC Measurements
Organizations: Department of Electrical and Computer Engineering, North Carolina State University, Raleigh, NC · Radio Systems Research, Nokia Bell Labs, Murray Hill, NJ, USA
Abstract
Wireless digital twins (DTs) rely on 3D environment models to predict radio propagation and support wireless-network decisions, yet these models are often initialized from imperfect 3D maps. Errors in building position, height, footprint, and orientation can therefore cause a high-fidelity propagation engine to simulate the wrong physical environment. In this paper, we study how a deployed wireless network can repair an existing DT using its own radio frequency (RF) measurements. In particular, we introduce Twin Residual Alignment and Calibration Engine (TRACE), a physics-grounded learning-based self-calibration framework that treats twin maintenance as residual alignment between the physical world and the current DT. Using the same sensing configuration as the physical measurements, TRACE ray-traces the current DT, coherently backprojects the measured and simulated RF onto a common world grid, and extracts the same local region around each building's current DT position. A multi-view corrector then fuses evidence across sensing nodes and neighboring buildings to predict a gated six-parameter correction per building, without relying on absolute layout or sensor ordering, and supports iterative correction through re-rendering. On 5,400 held-out samples from unseen simulated scenes at 28 GHz, TRACE reduces 3D position RMSE from 2.202 m to 0.302 m and yaw RMSE from 4.978° to 0.894°, outperforming ViT and U-Net baselines under changes in layout, building count, sensing-node count, and SNR. On measured 28 GHz RF data from the NIST outdoor courtyard, a model trained only on synthetic RF reduces mean planar wall-position error from 1.00 m to 7.8 cm, without measured-data fine-tuning or geometric labels. These results show that the discrepancy between measured and twin-rendered RF can serve as a learning signal for repairing a wireless DT.
Figures & tables
| Input blinding | Scene shuffling | |||||
| TRACE | only | only | Both blinded | Shuffled | Shuffled | |
| 3D RMSE [m] | 0.302 | 0.835 | 2.630 † | 2.460 † | 0.889 | 3.723 † |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| state of building in the current twin | |
| current twin state, buildings | |
| physical scene, expressed in the same parameterization | |
| target residual for building | |
| known acquisition configuration of sensing node | |
| measured and twin-rendered CSI at node |
| Input | 3D RMSE [m] |
|---|---|
| Amplitude only ( TRACE ) | 0.350 |
| Complex ( TRACE ) | 0.302 |
| Method | 2D [m] | Height [m] | 3D [m] |
|---|---|---|---|
| Uncorrected twin | 1.878 | 1.151 | 2.202 |
| TRACE , | 0.720 | 1.268 | 1.489 |
| TRACE , | 0.430 | 1.157 | 1.250 |
| TRACE , | 0.312 | 0.942 | 0.999 |
| TRACE , | 0.255 | 0.675 | 0.728 |
| TRACE , | 0.224 | 0.455 | 0.512 |
| Method | Type | 2D RMSE [m] | Error removed [%] |
| Corrupted twin | no correction | 1.878 | – |
| Phase correlation | training-free | 1.925 | |
| NCC | training-free | 1.802 | |
| ViT | learned | 1.483 | |
| U-Net | learned | 0.553 | |
| TRACE | learned | 0.206 |
| Component removed | Replaced by | Par. [M] | RMSE | Cost | Infl. [%] | Removed [%] |
| uncorrected twin | no correction | – | 2.202 | – | – | 0.0 |
| none | full TRACE | 3.98 | 0.302 | – | – | 86.2 |
| Cross-sensor attention | mean pooling only | 2.93 | 0.375 | 83.0 | ||
| Movement gate | raw residual applied | 3.98 | 0.332 | 84.9 | ||
| Sensor-pose FiLM | unconditioned encoder | 3.85 | 0.317 | 85.6 | ||
| Coherent residual stream | four-cue CueFusion | 3.91 | 0.315 | 85.7 |
| Corrector | ||||||
|---|---|---|---|---|---|---|
| U-Net on GeoCrops | 25.6 | 55.4 | 72.6 | 80.4 | 83.1 | 83.7 |
| TRACE | 52.8 | 72.6 | 81.3 | 84.9 | 86.2 | 86.6 |
| Model | Worst-case error removed [%] |
|---|---|
| TRACE | 79.6 |
| U-Net | 54.8 |
| ViT | 18.1 |
| Method | OOD-Layout | OOD-Error |
|---|---|---|
| Removed [%] | Removed [%] | |
| ViT | 23.59 | 10.60 |
| U-Net | 56.33 | 54.59 |
| TRACE , step 1 | 88.70 | 79.16 |
| TRACE , step 2 | – | 93.08 |
| Quantity | Value | Quantity | Value |
| Carrier frequency | Sensing nodes | 3, monostatic | |
| Bandwidth | Node height | (rooftop) | |
| TX array | Node downtilt | ||
| RX array | Node aim point | m | |
| Ray tracer | Sionna RT | Sensing-node building | 15\text{,}\mathrm{m}$$ |
| Stage 1: true-scene building parameters | |||
| Quantity | Value | Quantity | Value |
|---|---|---|---|
| Grid | px | Grid resolution | / px |
| Extent in | m | Extent in | m |
| Focus heights | 6 | Focus height values | m |
| GeoCrop size | px | GeoCrop extent † | square |
| Image scaling | per scene, shared | Scaling region | town crop, all heights |
| Quantity | Value | Quantity | Value |
|---|---|---|---|
| Token dimension | 256 | Tokens per crop | |
| Complex CNN stages | 4 | CNN channels (complex) | |
| Cross-attention | 1 layer, 4 heads | CueFusion projection | |
| Sensor fusion | 2 layers, 4 heads | Building attention | 3 layers, 4 heads |
| FiLM initialization | zero (identity at start) | Dropout | 0.2 |
| Parameters (full TRACE ) | 3,979,911 (3.98 M) | Absolute in graph | excluded |
| Quantity | Value | Quantity | Value |
| Optimization | |||
| Optimizer | AdamW | Learning rate | |
| Weight decay | Batch size | 128 | |
| Schedule | ReduceLROnPlateau | Warmup | 8 epochs |
| Max epochs | 200 | Early-stop patience | 20 epochs |
| Gradient clipping | 1.0 (max norm) | Mixed precision | enabled |
| Quantity | Value | Quantity | Value |
|---|---|---|---|
| Shared | |||
| Input channels † | Input size | ||
| Head trunk | Heads | , (6) | |
| Missing sensors | zero-filled channels | Sub-pixel correction | not applicable |
| U-Net | |||
| Base channels | 44 | Depth | 3 enc / 3 dec, bottleneck |