cs.LGOct 7, 2026

Multi-Agent Coordination via Support-Preserving Distillation

Authors: Sangmin Lee, Youngju Na, Chanmi Lee, Sung-eui Yoon

Organizations: School of Computing, KAIST

Abstract

Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples can be routed toward conflicting coordination modes. The teacher then produces samples between valid modes, and because the distillation loss regresses each local actor onto the conditional mean of the teacher's output given local input, this error is not absorbed but propagated to the student. To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training. We additionally study a shared-randomness variant that uses a shared noise component at execution to expose the residual gap intrinsic to strict-product execution. On controlled diagnostics and offline MARL benchmarks, MoSDOT improves endpoint quality and routing consistency, particularly on datasets exhibiting multimodal joint behavior.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CoDiMAD: Diffusion-Based Privileged Distillation for Communication-Free Multi-Robot Coordination

    Jul 10, 2026Jiyue Tao, Shunheng Xin, Tongsheng Shen +2Multi-Robot SystemsAware Heterogeneous Multi-Teacher Multimodal On-Policy Distillation

  2. CODA: Coordination via On-Policy Diffusion for Multi-Agent Offline Reinforcement Learning

    Apr 25, 2026Marcel Hedman, Kale-ab Abebe Tessera, Juan Claude Formanek +5Multi-Agent Reinforcement LearningDiffusion Policies

  3. MAS-OPD: On-Policy Distillation for Multi-agent Systems

    Sep 28, 2026Qiyong Zhong, Mao Zheng, Mingyang Song +5Aware Heterogeneous Multi-Teacher Multimodal On-Policy DistillationToken-Level Supervision