Existing architectures for LLM-based multi-agent systems (MAS) cannot reliably and efficiently solve multi-step tasks at scale: they struggle to support large numbers of agents and concurrent tasks, tolerate failures, govern agent interactions, and accommodate the diverse planning and execution patterns different tasks require. We present PANDA, a decentralized architecture that connects a large collective of heterogeneous, independently administered agents, letting them discover each other's capabilities and self-organize into small specialized teams per task. PANDA scales by decoupling collective communication from team communication, allowing agents to participate in multiple teams simultaneously, load-balancing tasks across the collective, and scheduling concurrent work within each agent. PANDA further separates the underlying architecture from the orchestration strategy, supporting three planning and execution patterns (star, chain, and mesh) that can be selected according to the structure and requirements of each task. PANDA detects infrastructure and orchestration failures and recovers affected tasks by dynamically replanning around failed components. Finally, to provide governance without a centralized service that would limit scalability, PANDA uses a web-of-trust model to constrain agent interactions to established trust relationships. We evaluate PANDA on the HotPotQA benchmark, demonstrating that it scales to thousands of agents, assembles teams in milliseconds, matches state-of-the-art accuracy at up to 8x the efficiency, and sustains 100% task completion under faults where existing systems fail.
Figures & tables
System
Scalable
Fault Tolerant
Flexible
Governance
Decentralized
Magentic-One ( Fourney et al., 2024 )
✗
✗
✗
✗
✗
Internet of Agents ( Chen et al., 2025 )
∼
✗
∼
✗
✗
Symphony ( Wang et al., 2025 )
∼
∼
✗
✗
∼
MACNET ( Qian et al., 2025 )
∼
✗
∼
✗
✓
AgentNet ( Yang et al., 2026 )
∼
∼
✗
✗
✓
panda (our solution)
✓
✓
✓
✓
✓
Table 1: Comparison of panda with existing multi-agent systems. ✓ = supported, ∼ = partially supported, ✗ = not supported. Flexible refers to flexible orchestration patterns.
Figure 1: Internal design of a panda agent. Team Communication handles incoming and outgoing messages for any teams the agent is a member of. The Agent Registry maintains state about other agents in the collective, including their capabilities and governance information. Collective Communication manages propagation and reception of system-wide messages such as agent joins and leaves. The Team Manager assembles teams and maintains state about currently active ones. It also possesses two subunits that handle the logic for the different planning and execution topologies and replanning when failures occur. The Infrastructure Failure Detector (FD) detects failures and relays them both internally and to the rest of the collective. The Agent Backend is our wrapper around any custom or existing LLM Agent that maintains stateful information relevant to panda about the agent. In turn, the agent must expose a list of capabilities C and implement a function f(c∈C,request)↦respose .
Figure 2: An example of each topology. Green indicates where input is passed to the team, and red denotes who provides the output. In the star, I/O flows through the orchestrator. In the chain, input enters the first agent and exits the last. In the mesh, I/O can begin and end at any teammate.
Figure 3: Team assembly times for number of capabilities c=5 , 25 , and 50 across redundancies r=1 , 2 , 3 . Rare refers to when only 5 agents with a given needed capability exist. At c=50 , 100 agents are not sufficient for staffing the needed capabilities.
Complete Knowledge
Partial Knowledge
System
EM (%)
F1 (%)
Total RT (s)
EM (%)
F1 (%)
Total RT (s)
Magentic-One
42.5
48.52
851.1
25.5
29.99
1181.7
AgentNet
67.5
77.33
840.4
43.5
51.37
1588.4
Internet of Agents
72.5
81.26
2182.1
62.0
68.07
3949.1
panda-star
66.5
74.11
464.5
45.5
54.60
594.9
panda-chain
64.5
75.54
462.5
47.0
53.18
452.9
Table 2: Performance on HotPotQA benchmark without failures.
Complete Knowledge
Partial Knowledge
System
Completion (%)
F1 (%)
Comp. RT (s)
Failed RT (s)
Completion (%)
F1 (%)
Comp. RT (s)
Failed RT (s)
Magentic-One
0.00
0.00
–
8.14
0.00
0.00
–
8.92
AgentNet
0.00
0.00
–
5.08
0.00
0.00
–
15.91
Internet of Agents
0.00
0.00
–
100.63
0.00
0.00
–
122.63
panda-star
100.0
72.33
28.04
–
100.0
68.10
39.42
–
panda-chain
100.0
68.23
27.29
–
100.0
60.43
32.23
–
Table 3: Performance on HotPotQA under infrastructure failures.
Complete Knowledge
Partial Knowledge
System
EM (%)
F1 (%)
Correct RT (s)
Incorrect RT (s)
EM (%)
F1 (%)
Correct RT (s)
Incorrect RT (s)
Magentic-One
45.00
52.31
28.24
59.58
6.50
6.83
18.01
40.07
AgentNet
1.00
1.56
47.90
24.57
22.50
27.51
72.90
50.59
Internet of Agents
24.50
27.34
131.40
192.78
11.00
12.51
57.87
278.40
panda-star
64.00
71.19
13.93
23.69
49.50
59.05
15.91
29.42
panda-chain
58.50
69.33
19.69
23.66
53.50
62.89
22.73
25.32
Table 4: Performance on HotPotQA under orchestration failures.
Figure 4: Average trusted recall in a collective where each agent requires two certificates to join with five genesis agents. Shaded regions show the IQR of trusted recall across agents.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Task & Team
T
User task
C
Required capabilities
A′
Team of agents
r(c)
Redundancy: agents recruited per capability
ϕ(c)
Pool of r(c) agents matched to capability c
Plan
Appendix
Table 5: Summary of notation.
Figure 5: Average trusted recall in a 1000 agent collective where each agent requires one certificate to join with one genesis agent. Shaded regions show the IQR of trusted recall across agents.
Figure 6: Average trusted recall in a 1000 agent collective where each agent requires three certificates to join with five genesis agents. Shaded regions show the IQR of trusted recall across agents.
Figure 7: Average trusted recall in a 2000 agent collective where each agent requires two certificates to join with five genesis agents. Shaded regions show the IQR of trusted recall across agents.