CollabFlow: Recursive Self-Improvement of Agent Collaboration
Organizations: The Chinese University of Hong Kong, Shenzhen · Fudan University · University of Oxford
Abstract
Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator level, topology-only learning keeps verbatim exchange that propagates errors, and reward maximization on a system's own outcomes concentrates on a few teams. To address these challenges, we propose CollabFlow, an RSI system of Learned Agent Collaboration: a trainable Collab-Director constructs teams of complete Agents, a frozen executor runs them, and each round's outcomes retrain the director. Within each round, the edges of a collaboration graph carry protocols of Evidence-Conditioned Communication: a receiver adopts a differing answer only when the sender's evidence is stronger by a margin, so the director learns who communicates and how. Across rounds, we further propose Collaborative Trajectory Balance (CTB), a flow-based objective that credits each team once across its construction orders and targets a reward-proportional distribution over teams, so several good teams stay in play. We also bound how far this self-generated target moves between rounds, which shrinks as records accumulate. On twelve datasets, CollabFlow outperforms all baselines and keeps improving across rounds. Code is available at https://anonymous.4open.science/r/CollabFlow-631E.
Figures & tables
| Baseline | SFT | GRPO | AFlow | Agent+RL | Skill evolution | Ours | ||||
| Dataset | Metric | Qwen3.5 | Qwen3.5 | Qwen3.5 | Qwen3.5 | FlowSteer | Evolving | SkillFlow | SkillRL | CollabFlow ( ) |
| (a) In-Distribution (IID) benchmarks | ||||||||||
| HotpotQA | EM | 46.25 | 47.34 | 49.53 | 53.75 | 53.28 | 60.16 | 60.78 | 58.28 | 63.91 (+17.7) |
| F1 | 65.59 | 67.32 | 67.80 | 69.97 | 71.38 | 72.95 | 78.37 | 75.71 | 79.72 (+14.1) | |
| TriviaQA | EM | 69.69 | 69.53 | 69.84 | 87.34 | 87.34 | 86.88 | 86.56 | 82.97 | 87.97 (+18.3) |
| F1 | 74.67 | 77.03 | 76.62 | 89.04 | 88.55 | 89.71 | 89.41 | 86.28 | 91.23 (+16.6) | |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Object / field | Meaning |
|---|---|
| Agent identity and role | Stable profile identifier, role description, and local objective. |
| Compatibility | Supported task families and capability labels. |
| Tool permissions | Permitted tools and whether an operation can modify protected environment state. |
| Visibility | Task evidence, messages, and observations visible to the Agent. |
| Execution mode | Stateless Agent, mutable Executor , or read-only Advisor . |
| Action | Structural effect | Required conditions |
|---|---|---|
| add_agent | Add to . | ; profile, permission, capacity, and budget checks pass. |
| add_edge | Add to . | Both endpoints have been selected; is supported for the task and endpoint modes; the edge is not already present; relation and budget checks pass. |
| bind_skill | Attach library skill to Agent . | ; the skill grants no tool permission and never overrides write permission. |
| set_output | Set output configuration to . | The output is not yet set; the configuration is supported, all referenced Agents exist, and its execution is compatible with task permissions and budget. |
| stop | Commit the terminal organization. | The team is nonempty, output is defined, and all task-specific terminal constraints are satisfied. |
| Protocol | Execution meaning |
|---|---|
| FinalOnly | Make a completed result available without rerunning the receiver. |
| OneWay | Send evidence or a proposal to the receiver, which may revise once under the configured resource bound. |
| Interactive | Allow a bounded sequence of mutual correction rounds where the task and runtime support it. |
| Stateful restriction | Internal Advisors provide bounded one-way guidance to the sole mutable Executor, and write permission stays with the Executor. |
| Benchmark | Suite | Task | Metric | Eval. items |
|---|---|---|---|---|
| HotpotQA | IID | Multi-hop question answering | EM / F1 | 128 |
| TriviaQA | IID | Factual question answering | EM / F1 | 128 |
| AIME 2026 | IID | Mathematical reasoning | Accuracy | 30 |
| MedXpertQA-Text | IID | Medical reasoning | Accuracy | 128 |
| ALFWorld | IID | Embodied task completion | Success rate | 128 |
| MBPP+ | IID | Function-level code generation | Pass@1 | 128 |
| Method | Role in comparison | Configuration |
|---|---|---|
| Direct Qwen3.5-9B | Shared-backbone reference | Checkpoint, tools, prompt, decoding and submission limits. |
| SFT | Task-level supervised baseline | Demonstration source, training split, trainable components, optimization budget. |
| GRPO | Task-level RL baseline | Policy target, reward, sample group, optimizer and rollout budget. |
| AFlow | Workflow-search baseline | Search budget, backend, candidate components and task adapter. |
| FlowSteer | Workflow-orchestration baseline | Policy/backend versions, action interface, reward and adapter changes. |
| Evolving orchestrator | Adaptive-organization baseline | Orchestrator policy, Agent pool, reward and adapter changes. |
| Row | Intervention |
|---|---|
| Single agent with tools | One tool-using Agent answers alone. |
| Parallel vote | Independent Agents answer; a majority vote selects the output. |
| Fixed chain | A fixed researcher–solver–verifier chain with verbatim messages. |
| Fully connected debate | Every Agent receives every other Agent’s full output each round. |
| Evolving orchestrator | The orchestrator of Dang et al. (2026) activates Agents in sequence. |
| Director training | The untrained Collab-Director builds the teams. |
| (a) In-Distribution (IID) benchmarks | ||||||||
| HotpotQA | TriviaQA | AIME 2026 | MedXpertQA | ALFWorld | MBPP+ | Avg. (IID) | ||
| Executor | EM | EM | Acc. | Acc. | SR | Pass@1 | Direct | CollabFlow |
| GPT-5.6-Luna | 55.47 (+33.59) | 88.28 (+35.16) | 93.33 (+16.67) | 100.00 (+62.50) | 89.06 (+16.41) | 92.97 (+5.47) | 58.22 | 86.52 (+28.30) |
| DeepSeek V4 | 58.59 (+28.91) | 90.63 (+14.84) | 83.33 (+33.33) | 100.00 (+21.09) | 81.25 (+65.63) | 89.84 (+3.91) | 55.99 | 83.94 (+27.95) |
| GLM | 65.63 (+38.28) | 87.50 (+53.13) | 56.67 (+13.33) | 100.00 (+10.16) | 96.09 (+7.03) | 84.38 (+19.53) | 58.13 | 81.71 (+23.58) |
| DeepSeek V4.1 | 64.06 (+24.22) | 89.06 (+23.44) | 83.33 (+13.33) | 100.00 (+18.75) | 94.53 (+36.72) | 88.28 (+7.03) | 65.96 | 86.55 (+20.58) |
| Mean per item | Spread of the total | ||||||
|---|---|---|---|---|---|---|---|
| Method | Total | Input | Output | P90 | Max | Std | CV |
| Single agent | 12,010.0 | 10,889.8 | 1,120.2 | 27,507.4 | 39,279 | 13,892.5 | 1.157 |
| Fixed team | 10,404.4 | 9,119.6 | 1,284.8 | 13,091.0 | 14,437 | 2,373.8 | 0.228 |
| Ensemble | 24,081.8 | 20,894.2 | 3,187.6 | 37,974.8 | 47,942 | 12,557.5 | 0.521 |
| SkillFlow | 10,451.6 | 9,396.0 | 1,055.6 | 16,639.4 | 18,039 | 5,063.7 | 0.484 |
| CollabFlow | 9,676.6 | 8,645.8 | 1,030.8 | 11,490.0 | 11,986 | 1,903.2 | 0.197 |