OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime
Organizations: Shanghai Jiao Tong University · University of Cambridge · Nanyang Technological University · The University of Hong Kong · Imperial College London · Peking University · Tencent
Abstract
Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents clear attribution of observed gains. To this end, we introduce OpenCollab, a multi-agent coding framework that provides a unified infrastructure for programmable collaboration and controllable runtime. Specifically, OpenCollab unifies organization design, enforces experimental control on a shared runtime, and tracks execution through fine-grained event streams. On this basis, we define Adherence to quantify whether the declared organization is actually realized. Our experiments reveal that agents collaborate very differently across configurations: changing any single dimension shifts Adherence, from 47.2% to as high as 97.2%. Furthermore, extensive agentic coding benchmarks show that a two-coder workflow built on OpenCollab establishes new SOTA performance compared to the mainstream harnesses such as Mini-SWE-agent, Codex CLI, and Claude Code, showing that a well-designed organization can outperform strong existing harnesses, while OpenCollab's single-agent configuration uses the fewest tokens across all evaluated suites. OpenCollab establishes a unified multi-agent infrastructure for easy programmable collaboration and controlled causal evaluation.
Figures & tables
| Held: can the factor be set explicitly? | Verified: can the run be checked? | ||||||
| Artifact | Model | Tools | Budget | Context | Topology | Realized | Compared |
| Multi Agent frameworks | |||||||
| AutoGen ( Wu et al., 2024 ) | ✓ | ✓ | – | ✓ | ✓ | ||
| AG2 ( ag2ai, 2024 ) | – | – | – | – | – | ||
| LangGraph ( LangChain, 2023 ) | – | – | – | – | |||
| Coding agents | |||||||
| Benchmark | Harness | Pass@1 (%) | Avg. tokens (M) | Avg. cost (\downarrow$ | Cache hit (%) |
|---|---|---|---|---|---|
| SWE-bench Pro | Mini-SWE-Agent | 61.66 | 4.16 | 0.90 | 36.65 |
| Codex CLI | 63.73 | 7.31 | 0.73 | 90.27 | |
| Claude Code | 58.03 | 9.38 | 2.54 | 12.50 | |
| OpenCollab (Base) | 63.21 | 3.89 | 0.41 | 88.03 | |
| OpenCollab (Duo) | 64.25 | 6.50 | 0.73 | 84.95 | |
| Terminal-Bench 2.1 | Mini-SWE-Agent | 76.40 | 2.97 | 0.45 | 67.36 |
| Dimension | Variant | Adherence (%) | 95% CI | Unverified (%) | Pass@1 (%) | Tokens (M) |
|---|---|---|---|---|---|---|
| Model | Qwen3.8-Flash | 47.2 | [30.4, 64.5] | 13.9 | 75.0 | 3.05 |
| DeepSeek-V4.1-Flash | 66.7 | [49.0, 81.4] | 22.2 | 72.2 | 3.65 | |
| GPT-5.6-Luna | 97.2 | [85.5, 99.9] | 2.8 | 61.1 | 5.79 | |
| Tools | All tools | 47.2 | [30.4, 64.5] | 13.9 | 75.0 | 3.05 |
| No edit tools | 86.1 | [70.5, 95.3] | 0.0 | 66.7 | 3.24 | |
| Read-only | 94.4 | [81.3, 99.3] | 2.8 | 52.8 | 4.24 |
| Configuration | Adherence (%) | Pass@1 (%) | ITT (%) | CACE (%) |
|---|---|---|---|---|
| Single | – | 69.4 | – | – |
| Open card (reference) | 47.2 [47.2, 61.1] | 75.0 | [ , ] | |
| Mandatory card | 91.7 [91.7, 100.0] | 72.2 | [ , ] | |
| Budget stated, 2M | 75.0 [75.0, 83.3] | 69.4 | [ , ] | |
| Read-only | 94.4 [94.4, 97.2] | 52.8 | [ , ] |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | Adherence (%) | Tasks / | ITT | Wald | Paired, | Paired, |
|---|---|---|---|---|---|---|
| Open card | 47.2 | 17 / 19 | ||||
| Budget stated, 2M | 75.0 | 27 / 9 | ||||
| Mandatory card | 91.7 | 33 / 3 | ||||
| Read-only | 94.4 | 34 / 2 |
| Code | Table 2 row | Card | Model | Entry tools | Topology |
|---|---|---|---|---|---|
| S | – | – | Qwen3.8-Flash | all | – |
| S ′ | – | – | Qwen3.8-Flash | all | – |
| R | reference | Open | Qwen3.8-Flash | all | Star |
| M1 | Model | Open | DeepSeek-V4.1-Flash | all | Star |
| M2 | Model | Open | GPT-5.6-Luna | all | Star |
| T1 | Tools | Open | Qwen3.8-Flash | no edit | Star |
| Card | Closing section |
|---|---|
| Open | What is yours to judge. Nothing above tells you what order to do things in, whom to talk to, or how much of the work to do yourself. Keeping all of it and handing parts of it over are both open to you, and neither one is what you are expected to do. Those are your calls, and you make them the way you would judge any piece of work: what the request actually needs, what is worth another agent’s attention, what it costs to describe a piece of it well enough to hand over, and what it costs to carry all of it in one budget and one conversation. |
| Optional | What this run asks of you. Implementation is the Coders’. There are two of them and neither can see what the other is doing, so ask them one at a time: send the first a message with message_agent describing the problem and let it work out its own fix, and when its answer is back, write the second its own brief and let it work out its own. What goes into that second brief is yours – what you learned from the first is yours to use or to leave out. Which of the two this run delivers is yours to choose. You may depart from this and do the implementation yourself; that is your call, and there are runs where it is the better one. Use team_status first if you want to see who is live. Everything else about how you work – what you read, how you diagnose, what you put in each message – is still yours. |
| Mandatory | What this run asks of you. This run is not asking you to decide how to divide the work. It is asking you to divide it a particular way, so do that. Implementation is the Coders’. There are two of them and neither can see what the other is doing, so ask them one at a time: send the first a message with message_agent describing the problem and let it work out its own fix, and when its answer is back, write the second its own brief and let it work out its own. What goes into that second brief is yours – what you learned from the first is yours to use or to leave out. Which of the two this run delivers is yours to choose. Use team_status first if you want to see who is live. Everything else about how you work – what you read, how you diagnose, what you put in each message – is still yours. |
| Benchmark | Both pass | Only Duo | Only Base | Both fail | |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 (89) | 68 | 6 | 3 | 12 | 0.51 |
| DeepSWE (113) | 58 | 21 | 5 | 29 | 0.0025 |
| Adherence | Pass@1 | |||||
|---|---|---|---|---|---|---|
| Code | % [95% CI] | Unverified % | vs reference | % | vs reference | vs S |
| S | – | – | – | 69.4 | – | – |
| S ′ | – | – | – | 66.7 | 2:3 (1.00) | – |
| R | 47.2 [30.4, 64.5] | 13.9 | – | 75.0 | – | 3:1 (0.63) |
| M1 | 66.7 [49.0, 81.4] | 22.2 | 13:6 (0.17) | 72.2 | 2:3 (1.00) | – |
| M2 | 97.2 [85.5, 99.9] | 2.8 | 19:1 ( 0.001) | 61.1 | 2:7 (0.18) | – |
| Arm | Definition | Differs from Single in | Realized |
|---|---|---|---|
| Single | one agent works the task to completion | — | trivially |
| Workflow | shared role cards; a script decides the order of work, what each role receives, and when to stop | organization, enforced | by code |
| Team | shared role cards; the model decides whether to delegate, what to send, and when to stop | organization, offered | measured, |
| Layer | Files | Lines | What it holds |
|---|---|---|---|
| Public surface | 9 | 894 | The client ( sdk/client.py ); the modules for tools, environments, team files and workflows |
| Bootstrap | 22 | 6,690 | Loading a team file ( team_config.py , Appendix C.3 ); resolving tool names ( tool_registry.py ); assembling a session, a team or a workflow run |
| Adapters | 72 | 16,171 | Model providers and the per-model capability table ( llm/ ); the tools, among them the patch tool, the shell and message_agent ( tools/ ); local, container and worktree workspaces; the sandbox policy ( safety.py ); the per-run event stream ( trace.py , Appendix C.5 ) |
| Application | 51 | 15,202 | The session loop and its limits ( session_run.py ); the Team controller ( application/scheduler.py and its scheduler_* and scheduler* modules); the Workflow controller ( workflow.py and its workflow_* modules); context compaction ( shaping/ ) |
| Domain | 13 | 1,820 | The ten-state session machine and its legal transitions ( session.py , Appendix C.1 ); the team topology ( team.py ); the table of sessions that the Team controller keeps, as plain data ( domain/scheduler.py ); agents, events and token estimates |
| Event type | What it adds to the common fields | What it supports |
|---|---|---|
| Model call | Token counts and latency of the step | Per-role spend; the organizational overhead of Appendix C.5 |
| Tool execution | Tool name, arguments and result | Who briefed whom; which edges were addressed. Refused messages and refused tool calls are written as separate events |
| Agent finish | Whether the role produced content | Whether a teammate worked, read together with its model calls |
| Session termination | The role’s termination reason (Table 14 ) | Which runs were cut off by the allowance after delegation began |
| Refused spawn | The reason for the refusal | That the roster cannot change within a run (Section 3.2 ); that a single-agent run cannot coordinate |
| Worktree change | Commit sha | Which role’s commits reached which worktree (Section 3.4 ) |
| State | Recorded reason | When it is recorded |
|---|---|---|
| DONE | completed ; submitted | The role gave a final answer, or called the submit tool. |
| STOPPED | budget exceeded | At PRECHECK , the role had already spent its allowance. |
| budget exhausted before model call | The next call, priced before it was sent, would not fit in what remained. | |
| budget exceeded after model call | The last call cost more than its estimate and took the role past its allowance. | |
| step limit reached | The session reached its step limit. | |
| loop block limit reached | The role kept repeating a tool call without progress. |
| Axis | Realized value, read from | Deviation |
|---|---|---|
| Participation | the teammates that spent tokens and produced at least one model output | a declared teammate did no work |
| Delegation | a teammate that spent tokens and produced at least one model output | no such teammate, where the assigned regime offers or requires delegation |
| Role boundary | the tool calls attributed to each agent | a call outside the agent’s assigned tool set is executed |
| Budget sharing | each role’s token draw against its own allowance | a role reached its allowance and no refusal was recorded |
| Information flow | message_agent calls in the event log, in every arm | a message travels along an edge the topology does not declare |
| Context policy | the compaction level applied on each turn | a level other than the assigned one fires |