Census-Based Population Autonomy For Distributed Robotic Teaming
Authors: Tyler M. Paine, Anastasia Bizyaeva, Michael R. Benjamin
Organizations: Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA · Woods Hole Oceanographic Institution Woods Hole, MA 02543, USA · Sibley School of Mechanical and Aerospace Engineering, Cornell University, Ithaca, NY 14853, USA
Collaborating teams of robots show promise due to their ability to complete missions more efficiently and with improved robustness, attributes that are particularly useful for systems operating in marine environments. A key issue is how to model, analyze, and design these multi-robot systems to realize the full benefits of collaboration, a challenging task since the domain of multi-robot autonomy encompasses both collective and individual behaviors. This paper introduces a layered model of multi-robot autonomy that uses the principle of census, or a weighted count of the inputs from neighbors, for collective decision-making about teaming, coupled with multi-objective behavior optimization for individual decision-making about actions. The census component is expressed as a nonlinear opinion dynamics model and the multi-objective behavior optimization is accomplished using interval programming. This model can be reduced to recover foundational algorithms in distributed optimization and control, while the full model enables new types of collective behaviors that are useful in real-world scenarios. To illustrate these points, a new method for distributed optimization of subgroup allocation is introduced where robots use a gradient descent algorithm to minimize portions of the cost functions that are locally known, while being influenced by the opinion states from neighbors to account for the unobserved costs. With this method the group can collectively use the information contained in the Hessian matrix of the total global cost. The utility of this model is experimentally validated in three categorically different experiments with fleets of autonomous surface vehicles: an adaptive sampling scenario, a high value unit protection scenario, and a competitive game of capture the flag.
Figures & tables
Fig. 1: Overview of states: poses x , opinions z , and network graph G .
Fig. 2: Influence connections in the nonlinear opinion dynamics model ( 2 ). Adapted from [ 2 ] .
Fig. 3: Overview of hierarchical census model adapted from [ 10 ] . Opinion dynamics is used for distributed team assignment (blue). Locally, multi-objective behavior optimization is used to determine the best reference input, and a controller is used to execute the action (yellow).
Fig. 4: Taxa of decentralized approaches represented by the model in this paper. The CBPA model can be reduced to recover many existing approaches.
TABLE I: High Value Unit Protection Mission Design (Experiment 1)
Fig. 5: Overview of high value unit (HVU) protection mission. The moving patrol area is centered on the HVU (in black) with a smaller loiter area closer to the HVU. Protector vehicles (in green) allocate themselves among the options of patrolling, loitering, or intercepting intruder vehicles (in red).
Fig. 6: Patrolling vehicle balances two objectives via behavior optimization as investigated in [ 50 ] .
Fig. 7: Experimental demonstration of HVU protection mission with a 16ft OPT WAM-V USV as the HVU (circled in pink), and seven Heron USVs (four circled in green in this image) as the protector vehicles. Demonstration took place on the Charles River near the MIT sailing pavilion.
Vehicle
Intrusion Number
1
2
3
4
5
Abe
118.5
26.3
72.3
131.3
104.5
Ben
81.9
78.2
95.1
95.0
51.0
Cal
107.9
111.6
89.6
84.9
73.5
Deb
91.6
112.1
114.0
106.1
84.8
Max
118.4
53.8
43.0
106.8
96.3
TABLE II: Intercept cost (Distance in meters) at time of notification of intruder. Allocated vehicles are highlighted in green.
Fig. 8: Battery voltage in each Heron during a portion of the HVU mission. Periods with no intruders are marked in yellow. Battery level traces colored by option the vehicle selected. Top Battery voltage varied during the mission and dropped when vehicles increased speed to chase intruders. Bottom On average vehicles with higher battery levels were allocated to patrol as designed.
Fig. 9: Computation time per iteration by a Raspberry Pi 4 on all Herons during the HVU mission. Top Average NOD iteration time remains low as neighbors increase Bottom Average Helm iteration remains below 0.25 during the mission where up to five collision avoidance behaviors were active. Iteration increases to 0.5 above five when all the USVs are within the configurable radius of 25 meters that occurred only during deployment or retrieval.
TABLE III: Capture-The-Flag Game Mission Design (Experiment 2)
Fig. 10: Overview of Aquatics game of capture-the-flag (CTF) [ 51 ] [ 52 ] . Two teams (red and blue) of three vehicles score by reaching the opposing team’s flag and returning it to their flag location. Vehicles defending their own team’s flag can “tag out” intruders once the intruder crosses onto their zone of the field.
Ours vs.
Comp.
Flag
Flag Grabs
Flag
Flag Captures
Opponent:
Type
Games
Win %
Tie %
Grabs
Allowed
Captures
Allowed
Default [ 51 ]
Simulated
200
89.0%
6.0%
3.1 (mean)
0.95 (mean)
1.9 (mean)
0.2 (mean)
Best Rule-Based [ 51 ]
Simulated
200
65.0%
11.5%
1.9 (mean)
1.0 (mean)
0.95 (mean)
0.35 (mean)
No NOD (ablation)
Simulated
200
50.6%
11.7%
1.9 (mean)
2.0 (mean)
0.8 (mean)
0.5 (mean)
FS1*
Field
1
N/A
N/A
3
0
2
0
FS2*
Field
1
N/A
N/A
6
0
2
0
TABLE IV: Game performance in simulation and field competitions
Fig. 11: Example of distributed team assignment and non-convex utility functions generated by the combination of intercept and avoid collision behaviors. The three blue-team vehicles collectively choose the option to defend, and the intercept behavior uses the Hungarian Algorithm to determine optimal assignments as shown by the intercept points. In the case of the lower two blue vehicles, the optimal individual decision for heading and speed is computed using a combination of the utility functions from the intercept and collision avoidance behaviors.
Fig. 12: Field competition of 3 vs 3 Aquaticus at US Milliary Academy West Point. Three vehicles on each team (red and blue) play a game of capture the flag. Approximate location of field is outlined in yellow, flags are marked with triangles, and vehicles are marked with circles.
Fig. 13: Sea Robotics Surveyor M1.8 USVs used at the US Military Academy West Point.
Fig. 14: Collaborative adaptive search and sample mission where the goal is to find and sample time-varying hot spots (yellow). Vehicles are dynamically allocated between cooperative searching (pink) or cooperative sampling options (red). Network connections are shown in black.
Fig. 15: Method to compute the estimate of the marginal improvement in reward ( 48 ) by selecting the option to explore.
Fig. 16: State of system during the adaptive seek and sample mission. Unknown regions of interest are in red, and the grid values range from blue (no interest found) to red (found region of interest). Sampling location are the small red circles. Black lines indicate network connections between vehicles. Four vehicles are exploring within their Voronoi cell and their computed path with green vertices is directed towards the unknown sample locations as designed. Four vehicles are allocating samples via a distributed solution to the TSP. Vehicle paths with yellow vertices are computed by neighboring vehicles as the most probable path that vehicle would select if the neighbor selected to explore, given the neighbor’s estimate of the grid.
Fig. 17: Results using VPG for η1 , η3 , and η4 starting from different initial values. The results from [ 57 ] give an ablation study where there is no collective option allocation between sampling and search.
Fig. 18: State of the MRS and simulated environment approximately 50 minutes into a two hour mission on the Charles River [ 10 ] . Simulated blooms appeared randomly within both zones X-ray and Yankee, and grew larger with time. A randomly generated storm periodically passed over the regions. At this time during the mission, 8 vehicles were searching and sampling in Zone X-ray. The communication range was artificially limited to 160 meters, and the restricted inter-vehicle communication is visualized in red. Both zones measured 300 meters by 350 meters.
Multi-robot systems must simultaneously optimize competing objectives while maintaining coordinated behavior. Existing multi-agent reinforcement learning approaches often rely on fixed or centralized coordination, which limits adaptability and violates distributed constraints. This work introduces the Coordination-Informed Multi-Objective Reinforcement Learning (CIMORL) framework, integrating a distributed weight prediction mechanism, a privileged expert training strategy, and theoretical guarantees for Pareto-optimal solutions. We present the base CIMORL method alongside two sampling-based variants, CIMORL-TS (Tree Search) and CIMORL-MPPI (MPPI), which leverage privileged global information during training to enable fully decentralized deployment. Experimental validation in cooperative and adversarial scenarios demonstrates a 21.2% hypervolume improvement and superior policy stability compared to state-of-the-art baselines. Real-world experiments with Crazyflie drones further validate the framework's robustness in resource allocation and multi-attacker multi-defend scenarios under partial observability.
Antonio Marino, Esteban Restrepo, Soon-jo Chung +2
University of Cambridge · CNRS, Univ Rennes, Inria, IRISA · Division of Engineering and Applied Science, California Institute of Technology
Collaborative robots are well-suited to maritime missions that benefit from coordination, such as the exploration of unknown reef structures, inspection of subsea infrastructure, or search-and-rescue operations. These missions typically provide sparse feedback signals for measuring progress and require adherence to safety and regulatory norms, turning a mission into a multi-objective optimization problem. Coevolutionary algorithms can process these sparse feedback signals to generate coordinated behaviors, and in some cases extend behaviors to multiple objectives. However, incorporating high-level team objectives with low-level compliance considerations on the fly to balance norm adherence with team performance remains elusive. This paper introduces a multi-objective framework that blends coevolved behaviors with compliance behaviors to achieve a balance between maximizing team progress and minimizing norm violations. The key insight is to decouple learning from compliance since operational norms are prescribed rather than discovered. We demonstrate that our framework achieves high team performance while avoiding collisions on a collaborative swimmer rescue mission with up to 8 vehicles in a hardware deployment, and 12 vehicles in simulation. The key contribution of this paper is Marine Multi-Objective Compliance-Integrated Coevolution (MMOCIC), a framework that blends team-wide optimization with established norms for real-world deployments of learning-based coordination.
Everardo Gonzalez, Tyler M. Paine, Manuel Agraz Vallejo +3
Collaborative Robotics And Intelligent Systems Institute, Oregon State University, Corvallis, OR, USA · MIT Marine Autonomy Lab, MIT, Cambridge, MA, USA
Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch through a doorway to assembling structures on a construction site. However, achieving such coordination in mobile multi-robot settings remains challenging: centralized methods conditioned on the combined observations of a team scale poorly with team size, and decentralized methods that train one policy per robot often require explicit alignment procedures or information sharing at inference time to overcome partial observability. Our key insight is that the visuomotor priors of pretrained vision-language-action (VLA) models should enable reactive, decentralized collaboration from each robot's local observations alone, without these inference-time assumptions. We propose CHORUS, a framework that adapts a single VLA backbone to control diverse, multi-robot teams. At inference time, each robot runs an independent copy of CHORUS, conditioned only on its own observations and a robot-identifying prompt. In real-world experiments including mobile tape measurement, library book handovers, and laundry basket lifting, CHORUS achieves a 64% point improvement over decentralized, from-scratch models, improves reactivity to teammate behavior by 40% points, and outperforms centralized baselines. Together, these results show that a shared VLA backbone is capable of achieving decentralized multi-robot collaboration, without per-robot policies or inter-robot communication at inference.