AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC
Organizations: Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Clear Water Bay, Hong Kong
Abstract
Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense 6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.
Figures & tables
| Tier | Method | Plan-TSR (%) | CSR (%) | Dependency-F1 (%) | Capability-F1 (%) | First-Try (%) | Avg. Attempts |
|---|---|---|---|---|---|---|---|
| Normal | Direct LLM | 60.0 | 100.0 | 88.1 | 82.2 | 60.0 | 1.000 |
| Tool-only AIMS | 72.9 | 98.6 | 87.0 | 77.4 | 52.9 | 1.286 | |
| AIMS | 100.0 | 100.0 | 100.0 | 100.0 | 78.6 | 1.229 | |
| Challenging | Direct LLM | 48.6 | 97.1 | 63.6 | 61.3 | 48.6 | 1.000 |
| Tool-only AIMS | 60.0 | 98.6 | 72.5 | 65.3 | 45.7 | 1.471 | |
| AIMS | 88.6 | 100.0 | 92.4 | 90.3 | 68.6 | 1.257 |