cs.ROOct 5, 2026

Demo: Vision-Language Model-Guided Online Calibration of an Electromagnetic Digital Twin

Authors: Zerui Kang, Yishen Lim, Zhouyou Gu, Seungnyun Kim, Seung-Woo Ko, Tony Q. S. Quek, Jihong Park

Organizations: Singapore University of Technology and Design Singapore, Singapore · Inha University Incheon, Republic of Korea

Abstract

An electromagnetic (EM) digital twin gives mobile robots wireless situational awareness but depends on material conductivities that change with the environment. Online calibration faces initialization sensitivity and measurement travel costs. We demonstrate a vision-language model (VLM)-guided framework using a Unitree G1 robot and NVIDIA Sionna, with two VLM calls: material classification maps visible materials through ITU-R P.2040 to conductivity priors for Sionna's gradient descent on accumulated received signal strength (RSS) measurements; waypoint planning selects the next measurement location online using residual RSS calibration error and image coverage. In a real indoor scenario, the framework achieves a normalized mean absolute conductivity error of 1.74×10−41.74\times10^{-4} within 20 m of travel; random initialization never converges, while random waypoints require over twice the travel.

Figures & tables

Explore similar work

Sep 26, 2026cs.AI

TRACE: Learning to Self-Calibrate Wireless Digital Twins from ISAC Measurements

Wireless digital twins (DTs) rely on 3D environment models to predict radio propagation and support wireless-network decisions, yet these models are often initialized from imperfect 3D maps. Errors in building position, height, footprint, and orientation can therefore cause a high-fidelity propagation engine to simulate the wrong physical environment. In this paper, we study how a deployed wireless network can repair an existing DT using its own radio frequency (RF) measurements. In particular, we introduce Twin Residual Alignment and Calibration Engine (TRACE), a physics-grounded learning-based self-calibration framework that treats twin maintenance as residual alignment between the physical world and the current DT. Using the same sensing configuration as the physical measurements, TRACE ray-traces the current DT, coherently backprojects the measured and simulated RF onto a common world grid, and extracts the same local region around each building's current DT position. A multi-view corrector then fuses evidence across sensing nodes and neighboring buildings to predict a gated six-parameter correction per building, without relying on absolute layout or sensor ordering, and supports iterative correction through re-rendering. On 5,400 held-out samples from unseen simulated scenes at 28 GHz, TRACE reduces 3D position RMSE from 2.202 m to 0.302 m and yaw RMSE from 4.978° to 0.894°, outperforming ViT and U-Net baselines under changes in layout, building count, sensing-node count, and SNR. On measured 28 GHz RF data from the NIST outdoor courtyard, a model trained only on synthetic RF reduces mean planar wall-position error from 1.00 m to 7.8 cm, without measured-data fine-tuning or geometric labels. These results show that the discrepancy between measured and twin-rendered RF can serve as a learning signal for repairing a wireless DT.
Aug 9, 2026cs.LG

Trajectory Design and Budgeted Querying for Digital Twin Calibration

Digital-twin calibration requires interaction data that is expensive to collect. We study two acquisition decisions: which trajectories to generate, and when to spend a limited budget on privileged parameter measurements. Our framework couples an excitation-oriented reinforcement learning controller, a recurrent parameter estimator with predictive uncertainty, and a budgeted query policy. In Pendulum, a Random Forest diagnostic recovers gravity only weakly from task-oriented trajectories and does not recover mass or length, while a GRU trained on excitation-oriented trajectories reaches a mean absolute error of 0.0066 with no queries. We then withdraw continuous oracle access partway through an episode, so that the twin must run on the estimator's output for the remainder. The estimator-plus-policy pipeline achieves a terminal error of 0.0092 under a three-query budget, against 0.2031 for an uncalibrated twin. In partially observable Waterworld, five controllers produce different observed error profiles across three hidden parameters, and an estimator trained on a five-controller mixture reaches online normalized errors of roughly 4-5%. These exploratory case studies are not controlled ablations, but they motivate treating trajectory design and query allocation as explicit design variables in data-scarce calibration.
Apr 22, 2026cs.RO

Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models

Recent advances in vision-language-action (VLA) models for robotics have highlighted the importance of reliable uncertainty quantification in sequential tasks. However, assessing and improving calibration in such settings remains mostly unexplored, especially when only partial trajectories are observed. In this work, we formulate sequential calibration for episodic tasks, where task-success confidence is produced along an episode, while success is determined at the end of it. We introduce a sequential extension of the Brier score and show that, for binary outcomes, its risk minimizer coincides with the VLA policy's value function. This connection bridges uncertainty calibration and reinforcement learning, enabling the use of temporal-difference (TD) value estimation as a principled calibration mechanism over time. We empirically show that TD calibration improves performance relative to the state-of-the-art on simulated and real-robot data. Interestingly, we show that when calibrated using TD, the VLA's single-step action probabilities can yield competitive uncertainty estimates, in contrast to recent findings that employed different calibration techniques.