MAS-OPD: On-Policy Distillation for Multi-agent Systems
Authors: Qiyong Zhong, Mao Zheng, Mingyang Song, Houcheng Jiang, Jiajie Su, Huwei Ji, Li Zhang, Junfeng Fang
Organizations: University of Science and Technology of China · Foundation Model Department, Tencent · Zhejiang University · National University of Singapore
Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.
Figures & tables
Figure 1: Motivation of MAS-OPD. Prompt-only MASs are costly and hard to customize, while MAS-RL suffers from sparse outcome rewards and cumbersome reward design. Extending OPD to MASs raises two challenges: building complementary role specialization while preserving shared knowledge, and turning cross-agent information into supervision for joint decisions and coordination.
Figure 2: Overview of MAS-OPD. MAS-OPD consists of three stages: (1) joint rollout , where agents interact with the environment to produce an on-policy joint trajectory; (2) teacher supervision , where RAS derives role-specific supervision by contrasting teacher scores across roles and PAC provides teacher-only privileged attribution for coordination; and (3) policy update , where the OPD and role-advantage signals are combined into token-level updates of the independent agent policies.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Domain
Role
Program run
Test input
Code
Coder
the Coder’s
the dataset’s golden unit tests
Tester
the Coder’s
the Tester’s own test case
Math
Tool-User
the Tool-User’s
—
Reasoner
—
—
Appendix
Table 3: What is executed on behalf of each role. A row is one execution the environment performs, not a program the role wrote: the only program run on the code domain is the Coder’s, and the Tester contributes the test input it is run on, so the two code rows are the same program on two different inputs. The Reasoner writes no program at all and states its answer in natural language, which is compared against the Tool-User’s printed answer by a symbolic verifier.
On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single-teacher capability ceiling: when the teacher errs, the student inherits the error. OPD also remains largely unexplored in agentic tasks, where per-step errors compound across long trajectories and destabilize training. We propose MAD-OPD (Multi-Agent Debate-driven On-Policy Distillation), which breaks this ceiling by recasting the distillation teacher as a deliberative collective of teachers that debate over the student's on-policy state; the debate produces an emergent collective intelligence that supplies token-level supervision, with each teacher's contribution weighted by its post-debate confidence. To extend OPD to agentic tasks, we also introduce On-Policy Agentic Distillation (OPAD), which adds step-level sampling to stabilize training under multi-step error compounding. We additionally derive a task-adaptive divergence principle, selecting JSD (Jensen-Shannon divergence) for agentic stability and reverse KL (Kullback-Leibler) divergence for code generation, and verify it both theoretically and empirically. Across six teacher-student configurations (Qwen3 and Qwen3.5; 1.7B-14B students, 8B-32B teachers) and five agentic and code benchmarks, MAD-OPD ranks first across all six configurations; on the 14B+8B→4B setting it lifts the agentic average by +2.4% and the code average by +3.7% over the stronger single-teacher OPD.
Jianze Wang, Ying Liu, Jinlong Chen +7
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology · Alibaba Group
Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost. We propose Orchestration Reward Modeling (OrchRM), a self-supervised framework for evaluating orchestration quality without human annotations. OrchRM leverages intermediate artifacts from multi-agent executions to construct win-lose pairs for Bradley-Terry reward model training. Unlike existing MAS test-time scaling and orchestrator training frameworks that rely on costly sub-agent rollouts, OrchRM operates directly at the orchestration level, enabling efficient and high-performing reward-guided orchestrator training and MAS test-time scaling. OrchRM improves training efficiency by up to 10x in token usage while improving MAS test-time scaling performance by up to 8% in accuracy. These gains consistently transfer across multiple domains, including mathematical reasoning, web-based question answering, and multi-hop reasoning, demonstrating orchestration-level reward modeling as a scalable direction for robust multi-agent orchestration. Code will be available at https://github.com/Wang-ML-Lab/OrchRM.
King Yeung Tsang, Zihao Zhao, Vishal Venkataramani +5
Large language model (LLM)-based Multi-agent systems (MAS) have shown promise in tackling complex collaborative tasks, where agents are typically orchestrated via role-specific prompts. While the quality of these prompts is pivotal, jointly optimizing them across interacting agents remains a non-trivial challenge, primarily due to the misalignment between local agent objectives and holistic system goals. To address this, we introduce MASPO, a novel framework designed to automatically and iteratively refine prompts across the entire system. A core innovation of MASPO is its joint evaluation mechanism, which assesses prompts not merely by their local validity, but by their capacity to facilitate downstream success for successor agents. This effectively bridges the gap between local interactions and global outcomes without relying on ground-truth labels. Furthermore, MASPO employs a data-driven evolutionary beam search to efficiently navigate the high-dimensional prompt space. Extensive empirical evaluations across 6 diverse tasks demonstrate that MASPO consistently outperforms state-of-the-art prompt optimization methods, achieving an average accuracy improvement of 2.9. We release our code at https://github.com/wangzx1219/MASPO.
Zhexuan Wang, Xuebo Liu, Li Wang +4
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China.