cs.LGOct 1, 2026

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

Authors: Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang

Organizations: Institute for Interdisciplinary Information Sciences, Tsinghua University · Engineering Systems and Design Pillar, Singapore University of Technology and Design

Abstract

Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CoFlow: Coordinated Few-Step Flow for Offline Multi-Agent Decision Making

    May 2, 2026Guowei Zou, Haitao Wang, Beiwen Zhang +2Flow Map

  2. CODA: Coordination via On-Policy Diffusion for Multi-Agent Offline Reinforcement Learning

    Apr 25, 2026Marcel Hedman, Kale-ab Abebe Tessera, Juan Claude Formanek +5Multi-Agent Reinforcement LearningDiffusion Policies

  3. Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents

    May 28, 2026Wenhao Li, Xiangfeng Wang, Bo JinMulti-Agent Reinforcement LearningModel-Based Reinforcement Learning