GAMBIT: Learning to Plan Continuous Multi-Robot Trajectories
Authors: Rishabh Jain, Akmaral Moldagalieva, Lorenzo Magnino, Michael Amir, Keisuke Okumura, Ajay Shankar, Wolfgang Hönig, Amanda Prorok
Organizations: University of Cambridge, UK · Technical University of Berlin, Germany · National Institute of Advanced Industrial Science and Technology (AIST), Japan
GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to improve overall team performance. Such self-sacrificial behaviours are difficult to capture with manually designed heuristics, particularly in dense, interaction-rich environments. Focusing on double-integrator continuous dynamics, this work studies how to learn such coordinated heuristics over motion primitives for multi-robot trajectory execution. Our framework, GAMBIT, first learns coordinated motion-primitive selection through imitation learning and subsequently fine-tunes the policy through reinforcement learning. We further introduce a safeguarded rollout mechanism with backup trajectories that guarantees collision-free execution at all times. Experiments demonstrate that GAMBIT substantially outperforms a range of baselines, including centralised motion planners and decentralised reactive planners, while exhibiting strong scalability. In particular, it coordinates over a thousand robots with planning latency below a few hundred milliseconds in continuous domains.
Figures & tables
Fig. 1: GAMBIT framework overview. We aim to learn a coordination-aware GNN policy that provides heuristic preferences over motion primitives. The framework consists of (a) imitation pretraining and (b) RL fine-tuning. The former step learns from demonstration trajectories by treating motion-primitive selection as a classification task. The latter uses multi-agent RL to maximise multi-robot navigation performance through our safeguarded rollout mechanism. The resulting policy, combined with the safeguarded rollout, enables collision-free execution while retaining high scalability.
Algorithm 1 Safeguarded rollout with PIBT and backups
Fig. 2: Main result. For each scenario, an example instance is shown at the top, where filled circles represent robots and unfilled circles connected by lines indicate their corresponding goals. Grey circles, rectangles, and cylinders represent obstacles. The remaining plots show, from top to bottom, CSR (complete success rate), SOC (sum-of-costs; normalised by lower bounds based on start–goal distances), and computation time to derive the entire trajectories. SOC and computation time are averaged over successful instances for each planner, with shaded regions indicating one standard deviation; to ensure meaningful comparisons, we report these metrics only when the corresponding success rate exceeds 50%. db-LaCAM’s SOC often exceeds the displayed range; we use an orange dot and arrow to indicate the clipped value to preserve readability of the remaining curves.
Fig. 3: Scalability assessment with up to 1,024 robots in the Forest scenario. The left panel shows a snapshot of the entire workspace at a certain timestep. Circles represent robots, with solid lines indicating their trajectories as tails. The upper-right panel depicts the trajectory evolution within a specific region, highlighting coordinated behaviours produced by GAMBIT ; e.g., the ⋆ robot temporarily moves away from its goal to let others pass. The lower-right panel presents individual success rate (ISR) evolution for n={256,512,1,024} ; i.e., the percentage of robots that have reached their goals, as a function of elapsed simulation time. We run the planners on 32 instances, with each curve representing the solution over a different instance. The bottom-right panel shows the distribution around the median inference time/step (and the corresponding loop frequency) for computing actions for the entire team.
Fig. 4: Snapshot from a GAMBIT deployment with eight multirotors planning at 10Hz . (left-inset) Speed profile (m/s) for the robots in this instance. (right-inset) Top-down view of the arena and trajectories.
As robots are increasingly deployed in groups and share workspaces to execute real-world tasks, planning their concurrent motions around complex manipulation skills becomes essential. These skills involve continuous physical execution and may exhibit stochastic behavior, resulting in variable execution times and uncertain continuous trajectories. Existing planners either limit execution to single-robot scenarios, rely on open-loop paths, or use post-hoc scheduling that prevents dynamic coordination. In this paper, we address this gap by integrating stochastic skills into sampling-based multi-robot planning by formulating the problem as a Markov Decision Process (MDP) over a multi-modal composite roadmap. For stochastic skills, solving the MDP yields a reactive policy that allows controllable robots to dynamically adapt their motions in response to other robots' execution of manipulation skills. By resolving skill uncertainty directly at planning time, this approach avoids the pessimism of conservative baselines and unlocks robust, dynamic multi-robot coordination. Code for the planners is available at https://www.vhartmann.com/stochastic-skills.
William Schnyder, Valentin N. Hartmann, Stelian Coros
Multi-robot trajectory planning is a fundamental problem in multi-robot coordination but remains computationally challenging due to its nonconvex, multimodal, and high-dimensional nature. This work builds upon D4orm, a dynamics-aware diffusion-denoising framework, and develops a family of planning architectures for diverse operational requirements. Unlike conventional numerical optimization methods, D4orm employs sampling-based optimization to generate solution trajectories through massively parallel sampling, leveraging modern computing architectures such as GPUs. Its diffusion-denoising structure iteratively optimizes \textit{deformations} to candidate control trajectories, providing an efficient and versatile paradigm for generating kinodynamically feasible and conflict-free trajectories. Using D4orm as the building block for advanced planners, we present a decoupled planner for improved scalability, an online receding-horizon planner with feedback control, and a distributed planner for resource-constrained settings. Evaluations with differential-drive and holonomic robots in 2D and 3D environments demonstrate that D4orm-based approaches find high-quality solutions faster and more reliably than other sampling-based optimization methods, such as MPPI, as well as a learned diffusion-model-based method. We further demonstrate zero-shot deployment on ten real quadrotors with obstacles, large-scale deconfliction with 100 simulated robots, and fully onboard distributed `lifelong' operation with six ground robots. Overall, these results establish diffusion denoising as a scalable and reliable framework for multi-robot coordination. Code and video: https://github.com/proroklab/d4orm
Yuhao Zhang, Keisuke Okumura, Ajay Shankar +1
Department of Computer Science and Technology, University of Cambridge, U.K. · National Institute of Advanced Industrial Science and Technology (AIST), Japan
Coordinating multiple robots in shared environments requires generating feasible trajectories for each agent while accounting for interactions among agents. Centralized planning approaches become difficult to scale as the number of robots increases, while decentralized approaches that allow each agent to plan independently do not inherently account for inter-agent interactions. This paper presents a framework for coordinated multi-robot motion planning that combines decentralized generative trajectory planning with multi-agent reinforcement learning (MARL)-based coordination. Each robot independently generates candidate trajectories using a diffusion model trained on single-agent motion data, leveraging the generative model's ability to produce feasible and diverse trajectories. To reduce conflicts between agents, a centralized value function trained via MARL guides the reverse diffusion process through gradient-based steering, enabling interaction-aware trajectory generation without centralized joint planning or retraining of the generative model. This guidance follows an exponential tilting formulation, in which the value function biases the denoising distribution toward trajectories with higher expected multi-agent return. The framework is evaluated in a simulated maze environment with four mobile robots. Experimental results show that the proposed value-guided diffusion planning reduces the inter-agent interference rate from 55.4% to 41.8%, demonstrating that coordination can be effectively achieved while preserving the scalability of decentralized trajectory generation. These results suggest that MARL-based value guidance can effectively introduce coordination into decentralized generative planners without requiring a fully joint multi-robot model.
Suk Ki Lee, Venkata Sai Deepak Mutta, Hyunwoong Ko
School of Manufacturing Systems and Networks, Arizona State University, Mesa, AZ · Michael W. Hall School of Mechanical Engineering, Mississippi State University, Starkville, MS