Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms
Organizations: Shanghai Academy of AI for Science (SAIS) · AI3 Institute, Fudan University · Yunnan University
Abstract
As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable local law that can be shared across interchangeable agents and adapt coordination as populations or interaction conditions change, without redefining a global organization. We introduce Waggle, a shared anonymous policy over bounded local views that jointly selects task actions, semantic communication, and local commitment updates. Repeated execution of the same law allows coordination to form, persist, and reorganize online without explicit roles or global topology. To learn this law across interchangeable agents and evolving coordination, we develop Swarm-Consistent Distillation (SCD), combining anonymous-orbit consistency with rollout-grounded prediction of the next local coordination field, with no added inference-time components. Across diverse coordination settings, the same learned law remains effective as populations and interaction budgets change, retains over 96% of substrate-specific oracle quality, and transfers without retraining; SCD further improves reorganization after counterevidence. Together, these results show that LLM-agent organization can emerge and adapt through repeated execution of a learned local law.
Figures & tables
| Silo-Bench | SwarmBench | Overall | |||||||||||
| Population scaling: SR / msg. | SR by difficulty | Raw scores : Pursuit / Foraging / Synchronization | Task-normalized | Avg. rank | |||||||||
| Method | Level I | Level II | Level III | ||||||||||
| Interface-matched controls | |||||||||||||
| NoComm | 10.7/0.00 | 9.8/0.00 | 11.7/0.00 | 20.0 | 10.1 | 5.0 | 1.71/ 1.31 /0.30 | 1.30/0.89/0.10 | 7.00/2.86/0.00 | 0.222 | 0.153 | 0.627 | 9.17 |
| NativeMsg | 74.7/9.00 | 67.0/9.00 | 14.3/9.00 | 25.2 | 13.0 | 4.7 | 1.50/1.03/3.40 | 0.80/0.53/2.00 | 0.99/0.82/0.80 | 0.348 | 0.194 | 0.167 | 8.46 |
| FixedSwarm | 77.6/9.00 | 69.7 /9.00 | 15.0/9.00 | 26.6 | 13.9 | 4.5 | 1.80/0.84/4.20 | 2.20/1.16/2.80 | 6.19/2.33/1.50 | 0.389 | 0.366 | 0.615 | 6.38 |
| Method | Task-wise soft score | Strict solved fraction | Scale ret. | Efficiency | Overall | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Coloring | Consensus | Leader- Election | Matching | Vertex- Cover | Overall | Msg./agent | Avg. rank | ||||
| Interface-matched controls | |||||||||||
| NoComm | 0.10 | 0.88 | 0.05 | 0.04 | 0.08 | 0.28 | 0.18 | 0.23 | 0.64 | 0.0 | 6.00 |
| NativeMsg | 0.34 | 0.90 | 0.55 | 0.25 | 0.18 | 0.54 | 0.35 | 0.44 | 0.65 | 28.4 | 4.93 |
| FixedSwarm | 0.44 | 0.91 | 0.72 | 0.34 | 0.35 | 0.64 | 0.46 | 0.55 | 0.72 | 15.8 | 2.29 |
| Shared-controller variants | |||||||||||
| Objective | Imitation | Training objective | Deployed structure | Closed-loop reorganization | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| NLL | Orbit-head acc. (%) | Train-head loss | Summary agree. (%) | Exact record equiv. (%) | Frozen-probe loss (norm.) | Recovery (%) | Stale AUC | Silo msg./agent | ||||
| All | Event | Decoded | Admitted | All | Event | |||||||
| Decision only | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] |
| Orbit | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] |
| Field | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] |
| Waggle | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] | [0pt] |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Pursuit | Foraging | |||
|---|---|---|---|---|
| Method | Score | Msg. | Score | Msg. |
| NativeMsg | 3.70 3.33 | 1742.5 531.9 | 2.90 2.18 | 341.2 325.4 |
| FixedSwarm | 6.10 4.09 | 509.8 148.5 | 3.70 1.42 | 80.5 73.5 |
| Waggle | 6.20 7.76 | 430.3 69.8 | 3.80 2.39 | 57.0 16.1 |
| Method | Success (%) | Msg. | kB |
|---|---|---|---|
| NativeMsg | 100.00 | 54.00 | 49.21 |
| FixedSwarm | 100.00 | 54.00 | 50.04 |
| Waggle | 97.92 | 41.63 | 34.47 |
| Success (%) | Msg. | kB | NativeMsg / FixedSwarm kB | |
|---|---|---|---|---|
| 4 | 100.00 | 27.27 | 16.68 | 21.85 / 22.42 |
| 8 | 95.83 | 55.98 | 52.25 | 76.58 / 77.66 |
| Element | Learned local law | SwarmBench fixed realization | Silo-Bench fixed realization |
|---|---|---|---|
| Task | Chooses a legal task action or waits | Egocentric local operation | Private-shard operation or response |
| Trace | Chooses content, incident target, and whether to relay | Contact-local cue with bounded per-hop relay | Local-port evidence frame with finite-lifetime relay |
| Commit and reopen | Chooses whether to synthesize, commit, challenge, or revise | Support/coherence readiness; contradictory cue | Native constructibility; incompatible partial evidence |
| Safeguards | Makes no schema, provenance, duplication, or expiry decision | Validates provenance, suppresses repeats, bounds relay, and expires cues | Validates provenance, suppresses repeats, bounds relay, and expires frames |
| Flocking | Transport | |||||
| Method | ||||||
| Interface-matched controls | ||||||
| NoComm | 5.40 | 5.00 | 4.30 | 0.00 | 0.00 | 0.00 |
| NativeMsg | 6.00 | 5.70 | 5.00 | 0.25 | 0.20 | 0.10 |
| FixedSwarm | 7.60 | 7.20 | 6.50 | 0.45 | 0.55 | 0.45 |
| Organizational baselines | ||||||
| Silo-Bench | SwarmBench | |||||
|---|---|---|---|---|---|---|
| Method | SR (%) | SR (%) | Scale retention | msg./agent | P/F task-norm. | P/F task-norm. |
| Waggle | 68.5 | 56.8 | 82.9% | 8.25 | 1.230 | 1.380 |
| FixedSwarm | 15.4 | 8.6 | 55.8% | 9.00 | 0.810 | 0.860 |
| SwarmSys | 27.1 | 17.8 | 65.7% | 9.00 | 1.100 | 1.150 |
| JSON valid | Schema valid | Executable | Mode macro-F1 | |
| Qwen3-4B + LoRA | 100.00 | 100.00 | 100.00 | 100.00 |
| Task | Overall writes | Local suppression | Context safeguards |
|---|---|---|---|
| Pursuit | Candidate ; executed | Repeats blocked ; refreshes | Unverified challenges blocked ; route relays silenced |
| Foraging | Candidate ; executed | Repeats blocked ; refreshes | Unverified challenges blocked ; route relays silenced |
| Mode | First matching condition |
|---|---|
| Challenge | Valid positive-conflict write |
| Synthesize | Commitment or response |
| Deposit | Fresh private-evidence write |
| Relay | Write matching an incident trace |
| Explore | Non-wait execution field with no higher label |
| Abstain | Legal wait with no write or commitment |
| Component | Value |
|---|---|
| Teacher / student | GPT-4o (temperature 0.2) / Qwen3-4B-Instruct-2507 |
| Initial policy | Common semantic-record warm start |
| Decision-task coverage | SwarmBench Pursuit/Foraging; Silo-Bench |
| Adaptation | LoRA rank 32, alpha 64, dropout 0.05 |
| Target modules | q/k/v/o and gate/up/down projections |
| Decision train / validation | 6,912 / 800 local views |
| Benchmark | Decision only quality | Waggle quality | quality | 95% lower bound | Waggle/ Decision only |
|---|---|---|---|---|---|
| Silo-Bench | |||||
| SwarmBench |
| Endpoint | Decision only | Waggle | Paired | Paired 95% CI |
|---|---|---|---|---|
| Summary agreement | 72.4% | 95.2% | pp | |
| Eventful frozen-probe loss | 1.000 | 0.608 | ||
| Recovery | 80.0% | 88.5% | pp | |
| Stale-field AUC | 22.768 | 19.237 | ||
| Post-conflict msg./agent | 7.157 | 6.737 |
| Objective | Deployed orbit agreement (%) | Eventful frozen-probe field loss | Silo-Bench success | Late-conflict recovery | Stale-field AUC | Post-conflict messages/agent |
|---|---|---|---|---|---|---|
| Decision only | ||||||
| Waggle | ||||||
| Change | pp | pp |
| Method | Native object | Benchmark integration | Retained mechanism | Implementation / tuning |
|---|---|---|---|---|
| GPTSwarm | Computational DAG of operation nodes and information-flow edges. | Map proposal-producing agents or worker operations to DAG nodes; route over legal contacts; translate benchmark actions or final answers at the boundary. | Edge-probability optimization and topological execution. | Clean-room; select a population-matched topology on disjoint development cases, then freeze it for formal evaluation. |
| G-Designer | Task-conditioned communication graph over agents and a virtual task node. | Map task information to node features; decode a topology over legal contacts; translate exchanged content and benchmark actions or final answers. | Graph encoder/decoder, virtual task node, and learned topology generation. | Pinned released code; published defaults or disjoint-development selection; frozen parameters with per-task topology decoding. |
| ARG-Designer | Autoregressively generated agent composition and communication topology. | Map task information and the available agent pool to the graph generator; execute the generated collaboration graph within the benchmark-legal interaction space. | Task-conditioned autoregressive generation of crew composition, roles, and communication links. | Pinned released code; published defaults; frozen generator with per-task graph generation. |
| AgentNet | Decentralized, dynamically evolving agent DAG. | Feed benchmark observations and outcomes into the native execution loop; routing evolves within the legal interaction space. | Forward/Split/Execute routing, local memory, and outcome-based edge updates. | Clean-room; fixed defaults across populations; within-execution graph updates retained and reset per episode. |
| DMoA | Sparse step-wise activation over an expert-agent pool. | Give the router benchmark-valid context and history; selected agents return benchmark actions or final answers through the benchmark boundary. | Context/history-conditioned recurrent routing and sparse activation. | Clean-room; published defaults or disjoint-development selection; frozen router with per-step activation retained. |
| SwarmSys | Explorer–Worker–Validator population with adaptive profiles. | Map benchmark events into the native role loop; profile matching allocates work; translate benchmark actions or final answers. | Specialized roles, profile–event matching, and validation-driven reinforcement. | Clean-room; fixed defaults across populations; online matching and validation feedback retained and reset per episode. |
| Endpoint | Comparator | Score | Waggle | 95% CI | ||
|---|---|---|---|---|---|---|
| Silo-Bench SR (%) | 16 | AgentDistill | 79.0 | 86.3 | ||
| 32 | FixedSwarm | 69.7 | 79.0 | |||
| 64 | AgentDistill | 43.7 | 58.2 | |||
| SwarmBench norm. | 16 | AgentDistill | 0.439 | 0.540 | ||
| 32 | AgentDistill | 0.425 | 0.546 | |||
| 64 | ARG-Designer | 0.880 | 1.041 |
| Method | SR (%) | LLM calls / episode | Input tokens (relative) | Output tokens (relative) | Msg./agent |
|---|---|---|---|---|---|
| Agent Distillation | 43.7 | 576 | 8.50 | ||
| DMoA | 21.5 | 806 | 3.60 | ||
| SwarmSys | 26.7 | 864 | 9.00 | ||
| Waggle | 58.2 | 577 | 8.03 |
| Variant | Overall | ||
|---|---|---|---|
| Waggle without priority | 0.61 | 0.39 | 0.50 |
| Waggle with priority | 0.89 | 0.71 | 0.80 |
| Method | Pursuit | Silo-Bench II-16 | Silo-Bench III-26 |
|---|---|---|---|
| Waggle (full refresh) | 4.33 / 1867.0 / 22.77 | 1.000 / 406.5 / 7.62 | 0.794 / 742.6 / 7.12 |
| NativeMsg | 1.40 / 1856.3 / 167.88 | 1.000 / 210.3 / 9.00 | 1.000 / 411.8 / 9.00 |
| FixedSwarm | 4.60 / 1860.0 / 29.36 | 1.000 / 210.3 / 9.00 | 1.000 / 411.8 / 9.00 |
| G-Designer | 3.40 / 1720.0 / 10.50 | 1.000 / 285.0 / 6.10 | 0.900 / 505.0 / 5.90 |
| DMoA | 3.65 / 1480.0 / 8.90 | 1.000 / 235.0 / 5.40 | 0.950 / 425.0 / 5.20 |
| AgentNet | 3.00 / 1859.9 / 8.38 | — | — |
| Task | Full refresh | Matched reuse | Input reduction |
|---|---|---|---|
| Silo-Bench II-16 | 1.000 / 406.5 / 7.62 | 1.000 / 239.3 / 7.58 | 41.1% |
| Silo-Bench III-26 | 0.794 / 742.6 / 7.12 | 0.957 / 420.6 / 7.27 | 43.4% |
| Deployment | SwarmBench P/F score / oracle retention | Silo-Bench SR / oracle retention | Deployed LoRA adapters |
|---|---|---|---|
| Joint Waggle | |||
| Exposure-matched specialist pair | |||
| Oracle specialist pair (6,912 per specialist) |
| Training regime | Swarm unique views | Silo unique views | Aggregate unique views | Optimizer updates | Deployed LoRA adapters |
|---|---|---|---|---|---|
| Joint Waggle | Matched | ||||
| Exposure-matched specialist pair | Matched † | ||||
| Oracle specialist pair (6,912 per specialist) | Matched |
| Comparator | SwarmBench P/F normalized score | Silo-Bench SR | Silo-Bench paired 95% CI | Reading |
|---|---|---|---|---|
| Exposure-matched specialist pair | — | Joint is higher on both | ||
| Oracle specialist pair | / retained |
| Deployment | Recovery | Re-synthesis latency | Relative stale AUC | Post-conflict messages/agent |
|---|---|---|---|---|
| Joint Waggle | ||||
| Exposure-matched specialist pair | ||||
| Oracle specialist pair |
| Decision source | Silo-Bench success | Silo-Bench messages/agent | SwarmBench P/F normalized score | Late-conflict recovery | Re-synthesis latency |
|---|---|---|---|---|---|
| GPT-4o local teacher | |||||
| Waggle | |||||
| Qwen3-4B zero-shot | |||||
| Fixed local heuristic | |||||
| Private-only | — | ||||
| Random legal record | — |
| Decision source | Recovery | Re-synthesis latency | Relative stale field AUC | Post-conflict messages/agent |
|---|---|---|---|---|
| GPT-4o local teacher | ||||
| Waggle | ||||
| Qwen3-4B zero-shot | ||||
| Fixed local heuristic | ||||
| Private-only | — | — |
| Population | Level I | Level II | Level III | Overall |
|---|---|---|---|---|
| 94.5 | 87.8 | 76.6 | 86.3 | |
| 91.8 | 80.7 | 64.5 | 79.0 | |
| 87.6 | 69.9 | 17.1 | 58.2 |
| Derived mode | Support | Predicted | F1 (%) |
|---|---|---|---|
| Explore | 200 | 200 | 100.00 |
| Deposit | 100 | 100 | 100.00 |
| Relay | 100 | 100 | 100.00 |
| Challenge | 200 | 200 | 100.00 |
| Synthesize | 100 | 100 | 100.00 |
| Abstain | 100 | 100 | 100.00 |
| Variant | Recovery | Rec. (pp) | Pooled Waggle signature |
|---|---|---|---|
| NoConflict | 0/40 | +90.0 | No explicit recovery in any group |
| NoDecay | 33/40 | +7.5 | 3.89 rounds faster; 86.0% less stale AUC |
| Irrev. Quorum | 0/40 | +90.0 | No explicit recovery in any group |
| Execution regime | Quality / sync | Recovery | Re-synthesis latency (round-equiv.) | Relative stale AUC | Msg./agent attempted/delivered |
|---|---|---|---|---|---|
| Synchronous | 1.000 | 36/40 (90.0%) | 4.5 | 8.2 / 8.2 | |
| Delay | 0.986 | 35/40 (87.5%) | 5.0 | 8.3 / 8.3 | |
| Delay | 0.969 | 34/40 (85.0%) | 5.5 | 8.4 / 8.4 | |
| 10% message drop | 0.951 | 33/40 (82.5%) | 5.8 | 8.4 / 7.6 | |
| 20% message drop | 0.905 | 30/40 (75.0%) | 6.7 | 8.5 / 6.8 | |
| 10% temporary churn | 0.925 | 32/40 (80.0%) | 6.1 | 8.6 / 8.1 |