Embodied Semantic Communication for Collective Autonomous Agents: A Tutorial on Representation, Wireless Delivery, and Closed-Loop Coordination
Authors: Yizheng Huang, Wensheng Lin, Lixin Li, Qinghe Du, Wenchi Cheng, Zhu Han
Organizations: School of Electronics and Information, Northwestern Polytechnical University, Xi’an, Shaanxi 710129, China · School of Information and Communications Engineering, Xi’an Jiaotong University, Xi’an 710049, China · School of Telecommunications Engineering, Xidian University, Xi’an 710071, China · Department of Electrical and Computer Engineering, University of Houston, Houston, TX 77004, USA
As autonomous systems and embodied intelligence enter the dynamic physical world, multi-agent collaboration calls for a paradigm shift in communication design. However, existing communication paradigms overlook that agents form action understanding from their own states, environmental observations, and collaboration relations through a process that evolves as a task unfolds. Consequently, reliable bit delivery, general semantic recovery, or single-task utility optimization alone cannot ensure that heterogeneous agents form coordinated actions compatible with their own conditions from shared information during task execution. To address this gap, this paper proposes embodied semantic communication (ESC) as a paradigm that transforms information transmission into action-oriented semantic interaction. Specifically, ESC characterizes how an explicit communication link can encapsulate multimodal perceptual states, intrinsic hardware capabilities, and collaborative intents into unified actionable semantic representations, thereby enabling heterogeneous receiving agents to parse, align, and ground them in local motor control. This paper clarifies the conceptual boundary, system characteristics, and environment-constrained technical pathways of ESC. It maps the underlying mathematical tools, including semantic information theory, world models, and multi-agent decision theory. Finally, this paper summarizes key open challenges, including measurable semantic reliability, ambiguity-triggered interaction under dynamic environments and tasks, and bandwidth-adaptive semantic transmission, outlining a roadmap for collective embodied networks.
Figures & tables
Fig. 1: Convergence of wireless communication and embodied intelligence toward embodied semantic communication. Wireless communication evolves from voice and messaging toward intelligent and semantics-aware connectivity, while embodied intelligence progresses from rule-based robotics to learning-enabled and foundation-model-driven systems. At their convergence, ESC organizes semantic formation, transmission and alignment, collective action, and physical feedback around embodied semantics.
TABLE I: Comparison of three communication paradigms.
Reference
Year
Primary Focus
Main Contributions
Lu et al. [ 23 ]
2024
General semantic communication
Reviews architectures, enabling technologies, evaluation metrics, applications, and open issues in semantic communication.
Qin et al. [ 24 ]
2024
AI-empowered wireless communications
Presents an overview of AI/ML-empowered wireless communications at the physical and lower MAC layers and AI/ML-enabled semantic communication systems.
Wu et al. [ 25 ]
2025
AI-enabled integrated sensing, communication, and computation
Reviews key technologies, system architectures, evaluation metrics, and AI integration for ISCC systems.
Chaccour et al. [ 26 ]
2025
Knowledge- and reasoning-driven semantic networks
Proposes an end-to-end framework for semantic representations, semantic languages, causal reasoning, knowledge accumulation, and reasoning-oriented performance measures.
Zhang et al. [ 27 ]
2025
Intellicise wireless networks
Presents a framework for intellicise wireless networks based on semantic communication.
Liang et al. [ 28 ]
2025
Generative-AI-driven semantic communication
Reviews transceiver design, semantic-effectiveness evaluation, knowledge management, network management, and applications of generative-AI-driven semantic communication.
TABLE II: Comparison of representative surveys.
Fig. 2: The structure of this paper.
Fig. 3: Closed-loop framework of embodied semantic communication. Heterogeneous agents encode local observations, body states, task intents, and collaboration needs into action-relevant embodied semantics and exchange them through explicit communication links, with optional assistance from network and human-supervisory nodes. Each receiver aligns the received semantics with its own state, task role, and local environment to update decisions and actions, while physical execution and environmental feedback continuously reshape subsequent semantic communication throughout the task cycle.
Fig. 4: Formation of action-relevant semantics from body state, environment, task goals, and interaction history.
Fig. 5: Real-world environmental constraints in embodied semantic communication, their resulting semantic failures, and corresponding technical pathways for ensuring reliable embodied semantics in collective tasks.
Fig. 6: Four modeling questions and their theoretical foundations for ESC.
Fig. 7: Illustrative comparison of three communication schemes for heterogeneous-agent collaborative tracking. Observation sharing transmits a target image or coordinate update; conventional semantic communication compresses the observation into a latent representation and reconstructs the target semantics at the receiver; ESC updates embodied semantics according to task progress and agent states, forms tracking-handoff semantics when the UAV’s view becomes blocked, and enables the quadruped to map the message into a tracking action using its local perception and dynamical capabilities. Execution feedback then updates the collective task state.
Collaborative embodied artificial intelligence (CEAI) enables multiple physical agents to perceive, reason, and act cooperatively in dynamic environments. Effective communication is essential for CEAI, yet CEAI agents must exchange not only large multimodal observations but also task-relevant insights, intents, and interactive information over long horizons. This article investigates token communication (TokCom) as a native intelligence interface for CEAI, in which tokens serve jointly as compact semantic carriers for communication and fundamental inference units for generative foundation models (GFMs). We first discuss how TokCom supports insight sharing, intent alignment, and interactive control among embodied agents. We then propose a TokCom-assisted CEAI framework driven by a task-adaptive communication protocol. Comprising a compact codebook, syntax rules, and contextual examples, this protocol guides GFM-based transceivers to distill messages into compact tokens and reconstruct them after wireless transmission. A case study on collaborative object transport demonstrates that the proposed TokCom framework substantially reduces the source payload bit consumption while preserving task efficiency and showing robustness under noisy channels. Finally, we outline future research directions.
Peng Yi, Ying-Chang Liang
National Key Laboratory of Wireless Communications, and also with the Center for Intelligent Networking and Communications (CINC), University of Electronic Science and Technology of China (UESTC), Chengdu 611731, China · Institute of Fundamental and Frontier Sciences and the Center for Intelligent Networking and Communications (CINC), University of Electronic Science and Technology of China (UESTC), Chengdu 611731, China
Embodied AI is increasingly becoming agentic, shifting robots from perception--control pipelines towards closed-loop systems that can retrieve context, deliberate during execution, monitor feedback, and refine future behavior. In parallel, robotics research has also moved from single-robot autonomy towards multi-robot systems, driven by the need for wider sensing, distributed action, heterogeneous capabilities, and fault tolerance. As AI agents move from single-agent use towards multi-agent collaboration, robotics faces a parallel challenge: robot teams must move beyond sharing maps, task assignments, and datasets towards sharing the state produced by embodied agent loops. This article explores Embodied Collective Intelligence (ECI), a future multi-robot paradigm in which a robot team accumulates and uses world context, task progress, and skill experience as shared resources. Specifically, we first review how embodied AI is becoming agentic and how multi-robot cooperation has evolved. We then present Embodied Collective Intelligence through Co-Perception, Co-Action, and Co-Evolution. Finally, we use an illustrative navigation study to examine one concrete component of the concept: shared world-memory inheritance. The study shows that a newly added robot can benefit from merged team memory, but it is not intended as a full evaluation of the ECI framework. Taken together, the review and conceptual framework motivate Embodied Collective Intelligence as a direction for embodied multi-agent intelligence, while the case study grounds one measurable part of the concept.
Yuxuan Yan, Yuanyuan Jia, Qianqian Yang
College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou, China
Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's evolving understanding of the world. When agents can only partially observe their surroundings, coordination without communication is provably hard, but communication can, in principle, bridge this gap by allowing agents to share observations and align their world models. In this work, we examine whether LLM-based embodied agents actually realize the ability to communicate. We extend PARTNR, a benchmark for collaborative household robotics, with a natural-language dialogue channel that enables two agents with partial observability to communicate during task execution. To evaluate whether dialogue leads to genuine world-model alignment rather than superficial coordination, we propose a framework for measuring world-model alignment defined over per-agent world graphs: observation convergence (do private world models align over time?), information novelty (do messages convey what the partner lacks?), and belief-sensitive messaging (do agents model what their partner knows?). Our experiments across three LLMs reveal that dialogue reduces action conflicts 40 to 83 percentage points but degrades task success relative to silent coordination. Using our metrics, we characterize the gap between superficial coordination and genuine world-model alignment, and identify where current models fall on this spectrum. Project Website: https://uiuc-conversational-ai-lab.github.io/partnr-dial-wmd/
Vardhan Dongre, Dilek Hakkani-Tür
Siebel School of Computing & Data Science University of Illinois Urbana-Champaign