Organizations: Institute for Interdisciplinary Information Sciences, Tsinghua University · Engineering Systems and Design Pillar, Singapore University of Technology and Design
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.
Figures & tables
Figure 1 : Superior performance and improved training efficiency of OMAF over 10 standard tasks.
Figure 2 : The CTDE framework of OMAF. Each agent generates a one-step flow action ai and path score surrogate ℓi through embedding attention (up), while jointly optimizes the flow policies (down) with a shared critic Qϕ(s,a) and synchronized policy loss L(θ) .
Figure 3 : Performance comparison of OMAF against representative online MARL methods HATD3 and HASAC, diffusion-based OMAD, and extensions of flow-based methods MAFlowRL and MAMFPO on the MPE and MAMuJoCo benchmarks. Curves show average episode return over training steps across 3 random seeds, with shaded regions denoting one standard deviation. OMAF achieves superior performance, validating one-step flow policies for online MARL.
Figure 4 : OMAF achieves superior performance across different metrics shown in the radar chart.
Figure 5 : Ablation studies of OMAF on the Ant 2×4 task.
Table 1 : Hyperparameters used across all tasks in the MPE and MAMuJoCo environments.
Task
Policy Learning Rate
Critic Learning Rate
Gradient Clip Norm
Softmax Temperature β
Cooperative Navigation 5 Agents
3×10−6
1×10−3
10000
1×10−1
Ant 2×4
3×10−4
3×10−3
10000
1×10−2
Ant 2×4 d
1×10−4
3×10−3
10000
1×10−2
Ant 4×2
3×10−5
3×10−3
10000
1×10−2
HalfCheetah 2×3
3×10−4
3×10−3
10
1×10−1
HalfCheetah 6×1
3×10−4
1×10−3
1
1×10−2
Appendix
Table 2 : Task-specific optimization hyperparameters for the MPE and MAMuJoCo environments.
Figure 7 : Analysis of OMAF design choices on Ant 2×4 . (a) Performance with different numbers of denoising steps over 1 million training steps, demonstrating the effectiveness of one-step denoising for online policy optimization. (b) Performance with different numbers of Transformer depth over 1 million training steps. (c) Performance under different softmax temperature settings.
Figure 8 : State coverage comparison on representative dimensions (23 and 13) at 250K steps. We visualize the state occupancy within the replay buffers for HATD3, HASAC, diffusion-based OMAD, flow-based MAFlowRL and MAMFPO, and OMAF. Colored regions indicate visited states. OMAF achieves the broadest coverage, where orange regions are uniquely explored by OMAF compared with the diffusion-based OMAD, demonstrating superior exploration.
Task
HATD3
HASAC
MAMFPO
MAFlowRL
OMAD
OMAF (Ours)
Cooperative Navigation N=10
−464.3±16.8
−471.3±14.0
−478.4±8.3
−467.2±23.4
−445.1±3.3
−429.4±8.3
Cooperative Navigation N=20
−1445.1±41.3
−1452.1±18.3
−1680.7±45.8
−1540.6±56.1
−1801.4±36.8
−1433.4±14.7
Appendix
Table 3 : Performance comparison on Cooperative Navigation with 10 and 20 agents.
Figure 9 : Visualization of learned diffusion policies across four distinct MAMuJoCo tasks. We display snapshots of the agents at timesteps t∈{1,100,250,500} with the instantaneous velocity, demonstrating the stable and coordinated behaviors achieved by our OMAF algorithm.
Figure 10 : Visualization of OMAF and other baseline algorithms in MAMuJoCo task CoupledHalfCheetah. We display snapshots of the agents at timesteps t∈{1,100,250,500} with the instantaneous velocity, demonstrating the stable and coordinated behaviors achieved by OMAF.
Generative models have emerged as a promising paradigm for offline multi-agent reinforcement learning (MARL), but existing approaches require many iterative sampling steps. Recent few-step acceleration methods either distill a joint teacher into independent students or apply averaged velocity fields independently to each agent. Unfortunately, these few-step approaches hurt inter-agent coordination. We show that the efficiency-coordination trade-off is not inherent: single-pass multi-agent generation can preserve coordination when the velocity field is natively joint-coupled. We propose Coordinated few-step Flow (CoFlow), an architecture that combines Coordinated Velocity Attention (CVA) with Adaptive Coordination Gating. A finite-difference consistency surrogate further replaces memory-prohibitive Jacobian-vector product backpropagation through the averaged velocity field with two stop-gradient forward passes. Across 60 configurations spanning MPE, MA-MuJoCo, and SMAC, CoFlow matches or surpasses Gaussian policies, value-based methods, transformer policies, diffusion models, and prior flow baselines on episodic return. Three independent coordination probes confirm that CoFlow's improvements arise from inter-agent coordination rather than per-agent capacity. A denoising-step sweep shows that single-pass inference suffices on every configuration. CoFlow reaches state-of-the-art coordination quality in 1-3 denoising steps under both centralized and decentralized execution. Project Page: https://guowei-zou.github.io/coflow/
Offline multi-agent reinforcement learning (MARL) enables policy learning from fixed datasets, but is prone to coordination failure: agents trained on static, off-policy data converge to suboptimal joint behaviours because they cannot co-adapt as their policies change. We introduce CODA (Coordination via On-Policy Diffusion for Multi-Agent Reinforcement Learning), a diffusion-based multi-agent trajectory generator for data augmentation that samples conditioned on the current joint policy, producing synthetic experience which reflects the evolving behaviours of the agents, thereby providing a mechanism for co-adaptation. We find that previous diffusion-based augmentation approaches are insufficient for fostering multi-agent coordination because they produce static augmented datasets that do not evolve as the current joint policy changes during training; CODA resolves this by more closely simulating on-policy learning and is a meaningful step toward coordinated behaviours in the offline setting. CODA is algorithm-agnostic and can be layered onto both model-free and model-based offline reinforcement learning pipelines as an augmentation module. Empirically, CODA not only resolves canonical coordination pathologies in continuous polynomial games but also delivers strong results on the more complex MaMuJoCo continuous-control benchmarks.
Marcel Hedman, Kale-ab Abebe Tessera, Juan Claude Formanek +5
Inephany Ltd · Department of Statistics, University of Oxford · The University of Edinburgh +2
Diffusion-based planning has achieved strong results in single-agent offline reinforcement learning, yet scaling to many-agent systems remains intractable due to the curse of dimensionality in the joint trajectory space. We introduce MF-Diffuser, a framework that lifts trajectory planning to the Wasserstein space of trajectory distributions, where the propagation of chaos ensures a small representative subset of agents captures the full population dynamics. Our approach features a value-weighted chaotic entropy objective that reconciles generative fidelity with return maximization, and a hierarchical coarse-to-fine strategy that progressively grows the agent population during denoising. We establish end-to-end suboptimality bounds with four interpretable terms, revealing that mean-field approximation error scales as O(H2/N) while offline distribution shift provably does not grow with population size N, and prove the generated policy is an approximate mean-field Nash equilibrium with explicit convergence guarantees. Experiments on three mean-field RL benchmarks -- spanning stage games, sequential dynamics, and adversarial team competition -- show MF-Diffuser achieves the best return in the majority of settings, with the largest gains on suboptimal offline data and at extreme scales (N≥103).