OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime
Authors: Chun-Wah Hsu, Kai Gong, Yu Wu, Xianhe Chen, Mengyang Liu, Jie Li, Hanyu Li, Zhixuan Liu, +7 more
Organizations: Shanghai Jiao Tong University · University of Cambridge · Nanyang Technological University · The University of Hong Kong · Imperial College London · Peking University · Tencent
Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents clear attribution of observed gains. To this end, we introduce OpenCollab, a multi-agent coding framework that provides a unified infrastructure for programmable collaboration and controllable runtime. Specifically, OpenCollab unifies organization design, enforces experimental control on a shared runtime, and tracks execution through fine-grained event streams. On this basis, we define Adherence to quantify whether the declared organization is actually realized. Our experiments reveal that agents collaborate very differently across configurations: changing any single dimension shifts Adherence, from 47.2% to as high as 97.2%. Furthermore, extensive agentic coding benchmarks show that a two-coder workflow built on OpenCollab establishes new SOTA performance compared to the mainstream harnesses such as Mini-SWE-agent, Codex CLI, and Claude Code, showing that a well-designed organization can outperform strong existing harnesses, while OpenCollab's single-agent configuration uses the fewest tokens across all evaluated suites. OpenCollab establishes a unified multi-agent infrastructure for easy programmable collaboration and controlled causal evaluation.
Figures & tables
Figure 1: Perspectives on multi-agent collaboration: (a) organization, (b) benefit and cost, (c) configured versus observed execution, and (d) collaboration auditing.
Held: can the factor be set explicitly?
Verified: can the run be checked?
Artifact
Model
Tools
Budget
Context
Topology
Realized
Compared
Multi Agent frameworks
AutoGen ( Wu et al., 2024 )
✓
✓
–
✓
✓
×
×
AG2 ( ag2ai, 2024 )
–
–
×
–
–
–
×
LangGraph ( LangChain, 2023 )
–
–
×
–
–
×
×
Coding agents
Table 1: Audit of seven controlled evaluation conditions. ✓: by the artifact’s own settings and records; –: only through the researcher’s code, or in part; × : not met.
Figure 2: OpenCollab overview. Single, Team, and Workflow controllers share a session runtime. Configuration and execution traces support Adherence and resource costs auditing.
Benchmark
Harness
Pass@1 (%) ↑
Avg. tokens (M) ↓
Avg. cost ()\downarrow$
Cache hit (%) ↑
SWE-bench Pro
Mini-SWE-Agent
61.66
4.16
0.90
36.65
Codex CLI
63.73
7.31
0.73
90.27
Claude Code
58.03
9.38
2.54
12.50
OpenCollab (Base)
63.21
3.89
0.41
88.03
OpenCollab (Duo)
64.25
6.50
0.73
84.95
Terminal-Bench 2.1
Mini-SWE-Agent
76.40
2.97
0.45
67.36
Table 2: Cross-harness evaluation on SWE-bench Pro, Terminal-Bench 2.1, and DeepSWE.
Dimension
Variant
Adherence (%) ↑
95% CI
Unverified (%)
Pass@1 (%) ↑
Tokens (M) ↓
Model
Qwen3.8-Flash
47.2
[30.4, 64.5]
13.9
75.0
3.05
DeepSeek-V4.1-Flash
66.7
[49.0, 81.4]
22.2
72.2
3.65
GPT-5.6-Luna
97.2
[85.5, 99.9]
2.8
61.1
5.79
Tools
All tools
47.2
[30.4, 64.5]
13.9
75.0
3.05
No edit tools
86.1
[70.5, 95.3]
0.0
66.7
3.24
Read-only
94.4
[81.3, 99.3]
2.8
52.8
4.24
Table 3: Single-dimension ablations relative to the reference team. Unverified runs are those terminated before collaboration; brackets report 95% Clopper–Pearson confidence intervals.
Configuration
Adherence (%) ↑
Pass@1 (%) ↑
ITT (%) ↑
CACE (%) ↑
Single
–
69.4
–
–
Open card (reference)
47.2 [47.2, 61.1]
75.0
+5.6
+11.8 [ +9.1 , +11.8 ]
Mandatory card
91.7 [91.7, 100.0]
72.2
+2.8
+3.0 [ +2.8 , +3.0 ]
Budget stated, 2M
75.0 [75.0, 83.3]
69.4
0.0
0.0 [ 0.0 , 0.0 ]
Read-only
94.4 [94.4, 97.2]
52.8
−16.7
−17.6 [ −17.6 , −17.1 ]
Table 4: Deployment contrasts and adherence-adjusted estimates across configurations on SWE-bench Pro. ITT is the Pass@1 difference from Single. Adherence is the share of runs with Ar=1 . CACE is ITT divided by Adherence. Brackets show sensitivity to the labeling rule for unverifiable verdicts and are not confidence intervals.
Figure 3: Two execution trajectories from Table 3 on a shared time axis. Lanes denote agents; blocks represent model calls colored by tool usage.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Adherence (%)
Tasks Ar=1 / 0
ITT
Wald
Paired, Ar=1
Paired, Ar=0
Open card
47.2
17 / 19
+5.6
+11.8
−5.9
+15.8
Budget stated, 2M
75.0
27 / 9
0.0
0.0
−3.7
+11.1
Mandatory card
91.7
33 / 3
+2.8
+3.0
+3.0
0.0
Read-only
94.4
34 / 2
−16.7
−17.6
−14.7
−50.0
Appendix
Table 6: Wald ratio and paired contrasts against the Single agent on SWE-bench Pro. Paired, Ar=1 is the mean Team-minus-Single difference over the tasks whose Team run is adherent, the direct estimate of Eq. ( 10 ) under the product coupling above; Paired, Ar=0 is the same difference over the remaining tasks, the estimate of δ0 . By construction ITT=αadh⋅(Paired,Ar=1)+(1−αadh)⋅(Paired,Ar=0) . Differences in percentage points.
Code
Table 2 row
Card
Model
Entry tools
Topology
S
–
–
Qwen3.8-Flash
all
–
S ′
–
–
Qwen3.8-Flash
all
–
R
reference
Open
Qwen3.8-Flash
all
Star
M1
Model
Open
DeepSeek-V4.1-Flash
all
Star
M2
Model
Open
GPT-5.6-Luna
all
Star
T1
Tools
Open
Qwen3.8-Flash
no edit
Star
Appendix
Table 7: The configurations of the adherence experiment. Each row but S ′ is a row of Table 3 , or the single agent it is compared with in Table 4 ; S ′ reruns S unchanged. Every agent has a 2M-token allowance except in B1 and B3. Each configuration runs the 36 tasks once. S and S ′ are the single agent; Star has an Adopter and two Coders, and both rings use the Ring roster (an analyst, a coder and a tester).
Card
Closing section
Open
What is yours to judge. Nothing above tells you what order to do things in, whom to talk to, or how much of the work to do yourself. Keeping all of it and handing parts of it over are both open to you, and neither one is what you are expected to do. Those are your calls, and you make them the way you would judge any piece of work: what the request actually needs, what is worth another agent’s attention, what it costs to describe a piece of it well enough to hand over, and what it costs to carry all of it in one budget and one conversation.
Optional
What this run asks of you. Implementation is the Coders’. There are two of them and neither can see what the other is doing, so ask them one at a time: send the first a message with message_agent describing the problem and let it work out its own fix, and when its answer is back, write the second its own brief and let it work out its own. What goes into that second brief is yours – what you learned from the first is yours to use or to leave out. Which of the two this run delivers is yours to choose. You may depart from this and do the implementation yourself; that is your call, and there are runs where it is the better one. Use team_status first if you want to see who is live. Everything else about how you work – what you read, how you diagnose, what you put in each message – is still yours.
Mandatory
What this run asks of you. This run is not asking you to decide how to divide the work. It is asking you to divide it a particular way, so do that. Implementation is the Coders’. There are two of them and neither can see what the other is doing, so ask them one at a time: send the first a message with message_agent describing the problem and let it work out its own fix, and when its answer is back, write the second its own brief and let it work out its own. What goes into that second brief is yours – what you learned from the first is yours to use or to leave out. Which of the two this run delivers is yours to choose. Use team_status first if you want to see who is live. Everything else about how you work – what you read, how you diagnose, what you put in each message – is still yours.
Appendix
Table 8: The closing section of the Adopter’s card, the only text that differs between the Open, Optional and Mandatory cards; the rest of the card is identical. Text the Optional and Mandatory cards share is in gray; what distinguishes each card is in black.
Figure 4: The four organizations. Left: Duo, a Workflow, OpenCollab (Duo) of Table 2 ; every arrow is a handoff that code makes on every run, and the dashed one runs only when the selection rule does not decide. The other three are the topologies of Table 3 , run as Teams: an arrow means the agent at its tail can message the agent at its head, and whether it does is the model’s choice.
Benchmark
Both pass
Only Duo
Only Base
Both fail
p
Terminal-Bench 2.1 (89)
68
6
3
12
0.51
DeepSWE (113)
58
21
5
29
0.0025
Appendix
Table 9: OpenCollab (Duo) against OpenCollab (Base), paired by task. p is the exact two-sided sign test on the two middle columns.
Adherence
Pass@1
Code
% [95% CI]
Unverified %
vs reference
%
vs reference
vs S
S
–
–
–
69.4
–
–
S ′
–
–
–
66.7
2:3 (1.00)
–
R
47.2 [30.4, 64.5]
13.9
–
75.0
–
3:1 (0.63)
M1
66.7 [49.0, 81.4]
22.2
13:6 (0.17)
72.2
2:3 (1.00)
–
M2
97.2 [85.5, 99.9]
2.8
19:1 ( < 0.001)
61.1
2:7 (0.18)
–
Appendix
Table 10: Adherence and Pass@1 of every configuration, with paired comparisons (only this configuration succeeded : only the other did, exact sign test p ). Adherence for a Star run counts both Coders; for a Ring run, the coder and the tester.
Figure 5: How runs end, as a share of each configuration’s 36 runs (Appendix B.1 ). Completed: the entry agent submitted. Token budget: the entry agent used up its allowance before submitting. Wall clock: the run reached 5,400 s. Failed: no result after the run was launched again. Configuration codes as in Table 7 .
Figure 6: Tokens per run. (a) Each dot is a run; the bar is the mean and the red tick is the run’s allowance (each agent’s allowance times the number of agents), with “mean / allowance” in millions at the right. (b) The mean split across agents, in percent. In a Ring the three agents are the analyst, the coder and the tester.
Figure 7: Wall clock per run. (a) Each dot is a run; the bar is the mean and the red tick is the 5,400 s limit, with “mean / limit” in minutes at the right. (b) The mean split across agents, where an agent’s time is the duration of its own model calls and tool calls; grey is the harness’s time between turns, a thin sliver at the right of each bar.
Figure 8: Share of tokens by agent in each tenth of a run, from its start to its last model call, averaged over the configuration’s runs. One panel per team configuration of Table 7 with Qwen3.8-Flash.
Figure 9: Read-only variant (the adopter can only read files and adopt a commit), resolved: two coders fix the task and the adopter adopts Coder A’s commit. A SWE-bench Pro task from the teleport repository.
Figure 12: The session state machine that every session runs. Solid edges are the checked transition table; a transition not in the table raises an error instead of being taken. Two escapes bypass the table from any state: an unhandled fault ends the session in ERROR , and a cancellation from outside the loop ends it in STOPPED . STOPPED is a controlled stop at a limit or a refusal, with the reason recorded (Table 14 ), not a fault. Each end state returns to IDLE only when a new turn starts.
Figure 13: Two controllers over one session runtime. The Team controller opens a session for every declared role, each with its own token allowance; the model decides whether to hand work on with message_agent , and a message along an undeclared edge is refused and recorded. The Workflow controller is a Python script that fixes the order of work and opens one session per step; a single-agent run is one session with no controller above it. All sessions run the same code, which checks every limit before a model call, acts in the role’s workspace, and writes to one ordered event stream per run (Appendix C.5 ).
Arm
Definition
Differs from Single in
Realized
Single
one agent works the task to completion
—
trivially
Workflow
shared role cards; a script decides the order of work, what each role receives, and when to stop
organization, enforced
by code
Team
shared role cards; the model decides whether to delegate, what to send, and when to stop
organization, offered
measured, αadh
Appendix
Table 11: The three arms. The last column says how a run is known to have realized its arm. Only in the Team is it a measurement of the model’s choice. In the Workflow the script, not the model, issues every handoff it reaches.
Figure 14: The runtime’s four layers as a clean architecture. Every import points inward, from bootstrap to adapters to application to domain; the application declares ports, interfaces that the adapters implement, so it never imports a model provider, a tool or a workspace directly. Figure 15 gives the files in each layer.
Layer
Files
Lines
What it holds
Public surface
9
894
The client ( sdk/client.py ); the modules for tools, environments, team files and workflows
Bootstrap
22
6,690
Loading a team file ( team_config.py , Appendix C.3 ); resolving tool names ( tool_registry.py ); assembling a session, a team or a workflow run
Adapters
72
16,171
Model providers and the per-model capability table ( llm/ ); the tools, among them the patch tool, the shell and message_agent ( tools/ ); local, container and worktree workspaces; the sandbox policy ( safety.py ); the per-run event stream ( trace.py , Appendix C.5 )
Application
51
15,202
The session loop and its limits ( session_run.py ); the Team controller ( application/scheduler.py and its scheduler_* and scheduler* modules); the Workflow controller ( workflow.py and its workflow_* modules); context compaction ( shaping/ )
Domain
13
1,820
The ten-state session machine and its legal transitions ( session.py , Appendix C.1 ); the team topology ( team.py ); the table of sessions that the Team controller keeps, as plain data ( domain/scheduler.py ); agents, events and token estimates
Appendix
Table 12: The runtime’s layers, top to bottom. Apart from five statements in the command-line interface (see text), no module imports from a layer above its own. Lines counts every line of the layer’s Python files.
Figure 15: The two code bases. The runtime’s layers run top to bottom, each with its file and line counts (Table 12 ), and a module imports only from its own layer or the ones below; the one exception is five imports from the command-line interface into the bootstrap layer (dashed), which the experiments do not use. The evaluator reaches the runtime only through its public surface.
Event type
What it adds to the common fields
What it supports
Model call
Token counts and latency of the step
Per-role spend; the organizational overhead of Appendix C.5
Tool execution
Tool name, arguments and result
Who briefed whom; which edges were addressed. Refused messages and refused tool calls are written as separate events
Agent finish
Whether the role produced content
Whether a teammate worked, read together with its model calls
Session termination
The role’s termination reason (Table 14 )
Which runs were cut off by the allowance after delegation began
Refused spawn
The reason for the refusal
That the roster cannot change within a run (Section 3.2 ); that a single-agent run cannot coordinate
Worktree change
Commit sha
Which role’s commits reached which worktree (Section 3.4 )
Appendix
Table 13: The seven event types. Model calls and agent finishes settle whether a teammate worked; tool executions of message_agent show who briefed whom; a session termination, read with the token counts of the model calls, identifies the runs whose decision the allowance cut short.
State
Recorded reason
When it is recorded
DONE
completed ; submitted
The role gave a final answer, or called the submit tool.
STOPPED
budget exceeded
At PRECHECK , the role had already spent its allowance.
budget exhausted before model call
The next call, priced before it was sent, would not fit in what remained.
budget exceeded after model call
The last call cost more than its estimate and took the role past its allowance.
step limit reached
The session reached its step limit.
loop block limit reached
The role kept repeating a tool call without progress.
Appendix
Table 14: The end states of a session (Figure 12 ) and the reasons recorded with them. Each reason is a string whose fixed prefix is shown; some continue with a count, such as the tokens used. The wall-clock limit is not among them: the evaluator stops a run that reaches it from outside the runtime and records that in the run’s metrics row.
Axis
Realized value, read from
Deviation
Participation
the teammates that spent tokens and produced at least one model output
a declared teammate did no work
Delegation
a teammate that spent tokens and produced at least one model output
no such teammate, where the assigned regime offers or requires delegation
Role boundary
the tool calls attributed to each agent
a call outside the agent’s assigned tool set is executed
Budget sharing
each role’s token draw against its own allowance
a role reached its allowance and no refusal was recorded
Information flow
message_agent calls in the event log, in every arm
a message travels along an edge the topology does not declare
Context policy
the compaction level applied on each turn
a level other than the assigned one fires
Appendix
Table 15: The six adherence axes. The assigned value is written before the run, and the realized value is read from the event log.
Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks, such as coding, through parallelization and context isolation. However, adding agents in practice introduces inter-agent communication overhead, which incurs extra cost and can sometimes offset the efficiency gains. We formalize multi-agent orchestration as a graph partitioning problem that captures the communication-to-computation trade-off: task decomposition can shorten critical-path computation, but cross-agent dependencies require costly context transfer. We instantiate this view in repository-level software engineering and present Cohesion-aware Coder (Co-Coder), which builds dependency graphs from static analysis, isolates structural hub files, partitions the graph via community detection, and executes the partition with a dependency-aware scheduler. Across 28 real-world tasks on DevEval and CodeProjectEval, Co-Coder advances the Pareto-frontier over sequential and file-based parallel baselines as well as Claude Code with Agent Teams, lifting pass rate by up to 14.0%, achieving up to a 2.10x wall-clock speedup, and reducing API cost by up to 35%, with the largest gains on the most dependency-dense projects. Co-coder demonstrates how cohesion-aware orchestration can make parallel coding agents both theoretically grounded and practically efficient, suggesting a broader design principle for multi-agent systems.
Xu Yang, Lunyiu Nie, Ethan Chandra +3
1The University of Texas at Austin · University of Oxford
Multi-agent vibe coding promises to accelerate software development, yet existing benchmarks rely on synthetic environments that ignore practical time and monetary costs, conflate reasoning with communication, and reward only superficial completion. We introduce multi-agent from-scratch evaluation benchmark, MSEval, evaluating multi-agent coding on real-world tasks. Grounded in 10 authentic, full-stack projects across 10 domains, MSEval scores performance using hierarchical requirements and deterministic rubrics. Its execution engine, LegoGent, tests 10 collaboration topologies where agents coordinate via periodic sync intervals and deploy through native CI/CD pipelines. Concurrently, the automated grader TAgent dynamically probes implementations to jointly measure functional success, latency, and prefix-cached token cost. Across 100 runs, MSEval reveals that organizational topology rivals model capability in shaping the speed--cost--quality trade-off. For identical tasks and models, varying the topology shifts scores by over 30 points and doubles wall-clock time. Structured pipelines converge fastest with the highest quality, whereas heavy managerial oversight degrades performance. Ultimately, MSEval establishes a rigorous, reproducible standard for measuring how multi-agent teams actually build software. The benchmark is released at https://github.com/robinren03/MSEval.
Recent advances in multi-agent systems have shown great potential for solving complex tasks. However, when multiple agents edit a shared codebase concurrently, their changes can silently conflict and inconsistent views lead to integration failures. Existing multi-agent systems address this through workspace isolation (e.g., one git worktree per agent), but this defers conflict resolution to a post-hoc merge step where recovery is expensive. In this paper, we propose STORM, i.e., STate-ORiented Management for multi-agent collaboration. Specifically, STORM manages agent states by mediating their interactions with the shared workspace, ensuring that each agent operates on a consistent view of the codebase and that conflicting edits are detected and resolved at write time. We evaluate STORM on Commit0 and PaperBench across multiple LLMs. STORM outperforms the git-worktree-based multi-agent baseline by +18.7 on Commit0-Lite and +1.4 on PaperBench, while achieving comparable or better cost efficiency. Combined with single-agent runs, STORM reaches highest scores of 87.6 and 78.2 on the two benchmarks respectively, suggesting that explicit state management is a more effective foundation for multi-agent collaboration than workspace isolation. STORM can also be plugged into any multi-agent system seamlessly.
Mengyang Liu, Taozhi Chen, Zhenhua Xu +2
Shanghai Jiaotong University · Cortices AI · Emory University +1