Organizations: Computer Vision and Aerial Robotics Group, Department of Artificial Intelligence, Universidad Polit´ecnica de Madrid, C. de los Ciruelos, Boadilla del Monte, 28660, Madrid, Spain. · Computer Vision and Aerial Robotics Group, Centre for Automation and Robotics (CAR), Universidad Polit´ecnica de Madrid (UPM-CSIC), Calle Jos´e Guti´rrez Abascal 2, Madrid, 28006, Madrid, Spain. · Automation and Robotics Research Group (ARG), Interdisciplinary Centre for Security, Reliability, and Trust (SnT), University of Luxembourg, L-1359, Luxembourg, Luxembourg. · Autonomous Systems Laboratory, ASLAB, Universidad Polit´ecnica de Madrid (UPM-CSIC), Calle Jos´e Guti´rrez Abascal 2, Madrid, 28006, Madrid, Spain.
As robotic teams tackle increasingly complex tasks in dynamic and unstructured environments, effective coordination requires agents to maintain accurate, aligned representations of their environment, teammates, and mission state. We argue that Mutual Awareness, Shared Situational Awareness, and Team Situational Awareness --- three concepts widely invoked in the multi-robot systems literature --- are not interchangeable: they form a containment hierarchy in which each type subsumes the previous in scope, and span a distribution spectrum from fully individualized understanding (MA) to fully uniform understanding (SSA), with TSA occupying a mixed position. We establish this through a two-dimensional taxonomy organized along the object of awareness and awareness distribution axes. Treating these terms as synonyms obscures the precise coordination requirements each imposes on a robotic system. A Search and Rescue case study with heterogeneous aerial and ground robots grounds each taxonomy position in concrete coordination requirements. These findings provide a conceptual foundation for principled specification, design, and comparison of collective awareness in multi-robot systems.
Figures & tables
Mutual Awareness
Source
Definition
Macmillan et al. (2004)
The extent to which team members are informed of other team members’ behaviors
Schuster and Jentsch (2011)
The level of understanding participants have of each other’s actions and contexts, which improves efficiency and performance in a simulator.
Rognin and Blanquart (2001)
Awareness about each other’s activities and about each other’s awareness.
Schmidt (1998)
The perception and understanding of member A on the activity, including intention, status, possible results and influence, of member B.
Villa et al. (2008)
If member A could observe member B’s activities while member B could not observe A, then A has a watching awareness. If member A knows member B is observing him/her, but A himself/herself could not watch B, then A has a watched awareness.
Table 1: Definitions of Mutual Awareness, Shared Situational Awareness, and Team Situational Awareness in collective contexts.
Figure 1: UGV A plans a trajectory in its local occupancy grid and sends the cell coordinates to its CA module, which transforms them into coordinates in a common reference frame. The CA module transmits this shared representation to UGV B, whose CA module transforms the Cartesian coordinates into cells in B’s local cost map and updates them accordingly. This coordinate transformation chain enables mutual awareness of planned paths despite heterogeneous internal representations.
Figure 2: A UAV detects a survivor in an image and obtains bounding box coordinates in the image frame. The UAV’s CA module transforms these coordinates into real coordinates in a common reference frame and transmits them to a UGV. The UGV’s CA module transforms the received coordinates into a cell in its local occupancy grid, enabling the UGV to integrate the survivor location into its spatial representation. This transformation chain allows heterogeneous agents to share situational awareness despite different sensing modalities and internal representations.
Figure 3: UAV C detects a safe landing area in an image captured by its camera and obtains bounding box coordinates in the image frame. C’s collective awareness module transforms these coordinates into real Cartesian coordinates in a common reference frame and transmits them to another UAV. The receiving UAV’s CA module transforms the shared coordinates into voxels in its local voxel map, enabling it to integrate the landing zone location into its spatial representation. This shared understanding of a critical environmental feature supports coordinated team operations and mission planning.
Type of Awareness
Distribution
Object of Awareness
Mutual Awareness
Distributed
Team member’s intentions, beliefs, goals, and plans.
Shared Situational Awareness
Shared
Essential elements for coordinated task execution. The intersection of each individual agent’s awareness requirements.
Team Situational Awareness
Mixed
Individualized awareness each agent must maintain to effectively perform their own functions. Overlaps with agents performing the same functions.
Table 2: Categorization of each term related to awareness in multi-agent systems according to the distribution and object of the awareness.
Figure 4: Venn diagram of the objects of mutual awareness, shared situational awareness, and role-specific elements. All of them combined are objects of team situational awareness.
Figure 5: Venn diagram of the objects of each kind of awareness for a system made up of ground and aerial robots operating in a search and rescue scenario. Specific objects of awareness for each role, combined with mutual awareness and shared situational awareness objects, form the whole set of objects of team situational awareness.
Embodied AI is increasingly becoming agentic, shifting robots from perception--control pipelines towards closed-loop systems that can retrieve context, deliberate during execution, monitor feedback, and refine future behavior. In parallel, robotics research has also moved from single-robot autonomy towards multi-robot systems, driven by the need for wider sensing, distributed action, heterogeneous capabilities, and fault tolerance. As AI agents move from single-agent use towards multi-agent collaboration, robotics faces a parallel challenge: robot teams must move beyond sharing maps, task assignments, and datasets towards sharing the state produced by embodied agent loops. This article explores Embodied Collective Intelligence (ECI), a future multi-robot paradigm in which a robot team accumulates and uses world context, task progress, and skill experience as shared resources. Specifically, we first review how embodied AI is becoming agentic and how multi-robot cooperation has evolved. We then present Embodied Collective Intelligence through Co-Perception, Co-Action, and Co-Evolution. Finally, we use an illustrative navigation study to examine one concrete component of the concept: shared world-memory inheritance. The study shows that a newly added robot can benefit from merged team memory, but it is not intended as a full evaluation of the ECI framework. Taken together, the review and conceptual framework motivate Embodied Collective Intelligence as a direction for embodied multi-agent intelligence, while the case study grounds one measurable part of the concept.
Yuxuan Yan, Yuanyuan Jia, Qianqian Yang
College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou, China
Real-world robots often operate in settings where objective priorities depend on the underlying context of operation. When the underlying context is unknown apriori, multiple robots may have to coordinate to gather informative observations to infer the context, since acting based on an incorrect context can lead to misaligned and unsafe behavior. Once the underlying true context is inferred, the robots optimize their task-specific objectives in the preference order induced by the context. We formalize this problem as a Multi-Robot Context-Uncertain Stochastic Shortest Path (MR-CUSSP), which captures context-relevant information at landmark states through joint observations. Our two-stage solution approach is composed of: (1) CIMOP (Coordinated Inference for Multi-Objective Planning) to compute plans that guide robots toward informative landmarks to efficiently infer the true context, and (2) LCBS (Lexicographic Conflict-Based Search) for collision-free multi-robot path planning with lexicographic objective preferences, induced by the context. We evaluate the algorithms using three simulated domains and demonstrate its practical applicability using five mobile robots in the salp domain setup.
Collaborative Robotics and Intelligent Systems (CoRIS) Institute, Oregon State University, Corvallis, OR 97331, USA · Khoury College of Computer Sciences, Northeastern University, Boston, MA 02115, USA
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
Hanchu Zhou, Dechen Gao, Hang Wang +5
University of California, Davis · Microsoft Research · Analog Devices