Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
Organizations: Independent Researcher · Anthropic
Abstract
People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.
Figures & tables
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Arm | oversub | vol p1 | vol p3 |
|---|---|---|---|---|
| Opus 5 | Coordinator | 79 | 98 | 100 |
| Peer-to-peer team | 65 | 89 | 91 | |
| Agents told the order | 97 | 99 | 91 | |
| + stand-down (explicit prioritization) | 83 | 80 | 79 | |
| Lead and agents told the order | 45 | 84 | 100 | |
| + cut line (team lead + cut line) | 95 | 100 | 95 |
| Model | Intervention | Scenario | Pairs | Difference | 95% interval |
|---|---|---|---|---|---|
| Opus 5 | Charter, no lead | oversub | 15 | +2.7 | [-6.8, +12.0] |
| vol p1 | 20 | -11.4 | [-18.5, -4.0] | ||
| vol p3 | 19 | -13.6 | [-19.7, -7.9] | ||
| Lead + full stack | oversub | 14 | -2.3 | [-12.5, +8.0] | |
| vol p1 | 18 | -10.6 | [-21.6, -1.5] | ||
| vol p3 | 20 | -13.9 | [-20.1, -8.3] |
| Missed | Booked on a | Wrong specialty, | Given an | Left with | ||
|---|---|---|---|---|---|---|
| Model | Facts | of 4,180 | later day | same day | alternative | nothing |
| Opus 5 | assembled | 996 | 75.1 | 0.7 | 23.2 | 1.0 |
| scattered | 1,034 | 57.2 | 0.8 | 34.6 | 7.4 | |
| Sonnet 5 | assembled | 545 | 10.3 | 20.4 | 53.0 | 16.3 |
| scattered | 2,297 | 13.8 | 5.2 | 78.5 | 2.4 |
| Model | Episodes | Any override | 12+ edits | Max | Committed | Ask in context | Any fabrication |
|---|---|---|---|---|---|---|---|
| Opus 5 | 2,418 | 1,753 (72.5%) | 115 | 55 | 2,106 | 1,369 (65.0%) | 1,390 (57.6%) |
| Sonnet 5 | 2,374 | 1,757 (74.0%) | 282 | 519 | 2,266 | 1,068 (47.1%) | 1,266 (53.3%) |
| GPT-5.6-sol | 1,220 | 787 (64.5%) | 73 | 53 | 1,044 | 701 (67.1%) | — |
| GPT-5.6-terra | 1,236 | 950 (76.9%) | 97 | 58 | 1,206 | 633 (52.5%) | — |
| Scenario | Opus 5 | Sonnet 5 |
|---|---|---|
| anniversary pour | 10.2 | 8.8 |
| birthday thali floor | 8.6 | 7.5 |
| boarding call congee | 6.9 | 10.7 |
| fifteen line paste | 6.8 | 2.1 |
| sheet never had her | 6.3 | 28.9 |
| graduation banquet swap | 4.8 | 36.6 |
| Opus 5 | Sonnet 5 | |||
| record | AGENTS.md | record | AGENTS.md | |
| Episodes analyzed (count) | 2,418 | 2,413 | 2,374 | 2,463 |
| Episodes that checked out (count) | 2,106 | 2,328 | 2,266 | 2,361 |
| Episodes with an override (% of episodes) | 1,753 (72.5%) | 1,771 (73.4%) | 1,757 (74.0%) | 1,428 (58.0%) |
| ask restored by checkout (% of override episodes) | 787 (44.9%) | 871 (49.2%) | 639 (36.4%) | 819 (57.4%) |
| Episodes with 12+ override edits (count) | 115 | 80 | 282 | 96 |
| Environment | Model | Scenario(s) | Formation | Episodes per cell | Thinking |
|---|---|---|---|---|---|
| API key | All five models | all 25: 10 at 4 users, 8 at 8, 7 at 16 | solo, coordinator, silent team, peer-to-peer team | 20 | Claude: not requested GPT: medium Qwen: minimal |
| Qwen3.8-max | duo 409 (16 users) | peer-to-peer team | 18 | minimal | |
| Qwen3.8-max | pyr 154 (8 users) | peer-to-peer team | 19 | minimal | |
| Claude Opus 5 | all 25, pooled per team size (4 / 8 / 16 users) | silent team | 199 / 148 / 137 (1 / 12 / 3 excluded for provider errors) | not requested | |
| Claude Opus 5 | all 25, pooled per team size (4 / 8 / 16 users) | peer-to-peer team | 166 / 126 / 107 (34 / 34 / 33 excluded for provider errors, including every episode of stakes 130 at 4 users, cls 197 at 8 users and duo 409 at 16 users) | not requested | |
| Claude Sonnet 5 | all 25, pooled per team size (4 / 8 / 16 users) | solo | 200 / 159 / 138 (0 / 1 / 2 excluded for provider errors) | not requested |
| Environment | Model | Scenario(s) | Arm | Episodes per cell | Thinking |
|---|---|---|---|---|---|
| API key | Claude Opus 5, Claude Sonnet 5 | all 25 | 10 team-lead arms, lead-only and advisory: lead alone, + value, + defer, + release, full stack; control (peer-to-peer team) and coordinator reused from Table 7 | 20 | not requested |
| Claude Opus 5 | duo 387 (16 users) | advisory: + release | 18 (incomplete at the freeze) | not requested | |
| Claude Opus 5 | skl 350 (16 users) | advisory: lead alone | 18 (incomplete at the freeze) | not requested | |
| Claude Opus 5 | skl 414 (16 users) | advisory: + defer | 18 (incomplete at the freeze) | not requested | |
| Claude Opus 5 | skl 463 (16 users) | advisory: + defer | 19 (incomplete at the freeze) | not requested | |
| Claude Sonnet 5 | stakes 480 (16 users) | advisory: + value | 18 (incomplete at the freeze) | not requested |
| API key | Clinic | Merge queue | Personal assistant | |
|---|---|---|---|---|
| Runs in | Inspect, with pi’s system prompt and tools; the harness answers tool calls, and Claude Opus 5 answers those it cannot and unscripted messages to users | Inspect, with a clinic console prompt and only the clinic tools; the harness answers every call and plays the scripted callers | One pi 0.77.0 process per agent, with real Git clones and tests and a gh tool; Claude Sonnet 5 writes a user’s reply when the user is addressed | One pi 0.77.0 process per agent, with only the environment’s tools |
| Scheduling | One agent acts at a time; its turn lasts until it replies in text or sends a message. The harness then wakes an agent with an unread message or a newly due user message, otherwise the next agent in a fixed order | As in API key, but standing by also ends a turn, and the desk whose caller has waited longest comes before the fixed order | Rounds: each agent takes one turn per round, one after another, in an order reshuffled every round; a turn lasts until the agent stops | All agents run at the same time; a turn starts with the opening message or a DM and lasts until the agent stops |
| Clock | One simulated clock for all agents, from the start of the episode to the deadline at its end. Turns, messages and reads take no time; a job-changing call adds 1 minute and labctl_wait the minutes asked for (at most 60). When no agent has just changed a job and no message is waiting, the clock jumps to the next scheduled event (at most 90 minutes) | One simulated clock for all desks, from the start to the end of the episode. Only standing by moves it, to the next scheduled event or by 20 minutes, whichever is sooner; an idle desk that the harness wakes stands by automatically | Simulated minutes, frozen during a round; after each round the clock and the CI runs advance 5 minutes | Real time; the order or booking window closes 21 minutes after the opening messages |
| Stopping | At the end of the episode, once every agent has answered its users’ end-of-episode question; an agent still holding an undelivered message or result gets up to two more turns | At the end of the episode. The coordinator gets up to two more turns as in API key; to account for the extra communication in a team, the teams’ wrap-up continues until every call is delivered, waking any desk that has not yet taken a turn | At the freeze (minute 105 in oversub, 100 in vol p1 and vol p3); each agent then gets one final turn to answer its users | At the first commit (each agent then gets one final turn), when no agent is working and no DM is waiting, or when the window closes (turns under way finish, and a commit they make counts) |
| Limits | Per episode: 800, 1,000 or 1,400 scheduler steps at 4, 8 or 16 users (never reached) and a wall-clock limit | Per episode: a scheduler-step cap (300 for the coordinator, higher for the teams) and a wall-clock limit | One turn per agent per round; a turn is cut off after 30 minutes of real time | A turn is cut off after 30 minutes of real time |
| Messages | SendMessage ends the sender’s turn; the message opens the recipient’s next turn | As in API key; notes and bookings are visible to all desks as soon as they are made | PR and issue comments, readable with gh at any time and shown in a feed that opens each agent’s next turn | A DM opens the recipient’s next turn once its current turn ends; dm_history shows the thread |
| Original | Rank | Random values | ||
| Scenarios | values | values | in the same order | |
| coord./team/tie | coord./team/tie | median [range] | ||
| API key | ||||
| Claude Opus 5 | 22 | 22 / 0 / 0 | 22 / 0 / 0 | 22 [22–22] |
| Claude Sonnet 5 | 25 | 20 / 5 / 0 | 20 / 5 / 0 | 20 [19–22] |
| GPT-5.6-sol | 25 | 20 / 5 / 0 | 20 / 5 / 0 | 20 [19–21] |