Nontrivial dynamics can emerge in large language model (LLM)-based multi-agent systems, and preliminary evidence exists that formalisms from statistical mechanics can be effective at modeling and predicting such behaviors. In parallel, designing multi-agent communication topology for optimal task-solving is an active research question. In this paper, we focus on predicting the success of multi-agent search tasks using the formalism of absorbing state phase transitions. We first taxonomize search tasks into four types, informed by classical results in combinatorial search. We then theoretically derive a critical communication degree dc, the minimum number of agents each agent can communicate with, above which incorrect hypotheses do not proliferate uncontrollably and the search enters the solved state. Finally, we evaluate frontier LLM-based multi-agent systems on real-world search and discovery tasks, software configuration debugging and physical mechanism discovery, and find that agreement with theory is mixed. LLM agents may not communicate with their neighbors and can develop strategies that are individually beneficial but limits the benefits of collaboration.
Figures & tables
Figure 1: Schematic of our multi-agent search setup. Agents begin with candidates each containing a single hypothesis, or literal ℓ (in this example, the literal asserts the answer to question q1 is 1). This candidate branches into b new candidates that add each a new hypothesis. In one search round, an agent (1) selects m of its active candidates to be verified; (2) runs the verification mechanism on the selected candidates to test whether they are consistent with the search target; (3) receives information about verification results from neighboring agents; and (4) the agent prunes incorrect candidates based on all available evidence. The next round begins by branching b new candidates from every surviving candidate of the current round.
Figure 2: Our search task taxonomy. Orange marks the target c⋆ , and red dashes mark what refuting a single literal can remove. The classes differ in how literals can be combined, and that difference determines evidence reusability and branching behavior.
Policy
Prompted objective
Communication rule
Team-Auto
Team performance
Automatically transmit the true verdict of every experiment to d neighbors (same as for deterministic agents).
Team-Free
Team performance
The agent chooses which of d neighbors to communicate with, what to report (deception is allowed), plus optional text.
Indiv-Free
Individual performance
Same as C1.
Table 1: LLM communication policies tested.
Figure 3: Deterministic agents recover theoretical dc prediction across all tasks. Filled points represent task configurations with multi-valued literals ( ∣Vi∣>2 ), such as in the TCAS task. Open points represent ∣Vi∣=2 , such as in the physical mechanism discovery task where each symbolic term either is or is not in the true equation.
Synthetic
TCAS
Physics
dctheory=6.10
dctheory=10.90
dctheory=6.10
Model
Team- Auto
Team- Free
Indiv- Free
Team- Auto
Team- Free
Indiv- Free
Team- Auto
Team- Free
Indiv- Free
Qwen3.5-4B
6.55 6.2–6.9
12.82 12.2–13.4
13.47 11.6–16.0
12.79 11.4–13.7
17.69 16.5–25.1
20.13 16.8–24.8
8.77 7.7–10.0
13.94 12.2–14.6
14.31 11.5–16.9
Qwen3.5-9B
6.64 6.2–7.0
16.01 13.8–17.4
16.06 14.1–18.2
12.90 10.9–15.2
29.15 23.6–37.3
27.31 24.8–29.6
8.34 7.7–8.8
17.07 15.0–18.9
16.77 14.2–18.5
GPT-5.4 Nano
6.39 5.8–7.1
7.87 7.0–9.0
7.29 6.8–7.8
10.52 8.7–11.6
12.88 9.8–14.5
14.41 9.7–16.1
10.95 9.8–12.0
13.04 11.3–16.7
12.79 11.8–13.4
GPT-6 Luna
6.88 6.4–7.3
7.06 6.8–7.3
6.46 6.0–7.0
10.05 9.6–11.2
13.71 10.0–18.0
13.13 9.3–14.9
8.17 7.3–9.0
8.48 7.9–9.2
7.88 7.4–8.8
Table 2: Empirical critical communication degrees dc for four LLM families. Estimates are obtained from pooled incorrect-lineage reproduction counts over six task seeds; error bars denote 95% task-level bootstrap confidence intervals. Lowest dc^ values are bolded, and the theoretical prediction for dc for each task is shown below each task name for comparison.
Figure 4: Higher reasoning effort on GPT-6-luna almost always raises the critical communication degree dc across all policies and tasks.
Figure 5: Changing the sampling temperature can produce very different dc for certain models and tasks. We test T=0.2 (filled point) against the default temperature setting T=1 (open circle) on all policies and tasks.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Synthetic hypothesis search graph. Each symbol ℓj,a denotes the atomic assignment that component j takes value a . Each refinement adds one previously unassigned component–value pair. Different assignment orders can reach the same hypothesis node, and experimentally refuted assignments prune all hypotheses that contain them.
Figure 7: TCAS configuration-fault hypothesis graph. Each child adds one parameter–value assignment to its parent. Different assignment orders that produce the same partial configuration correspond to the same hypothesis node. Blank nodes and ellipses indicate additional hypotheses omitted for clarity.
Figure 8: Physical-mechanism hypothesis graph. Each child adds one atomic symbolic component to its parent. Different addition orders can reach the same mechanism. Blank nodes and ellipses indicate additional hypotheses omitted for clarity.
Figure 9: Empirical success probability Psucc(d) as a function of communication degree d for representative horizons T=3,4,5 in the synthetic benchmark. Each point is the observed success fraction over eight independent task seeds, where an episode is successful if at least one agent recovers the ground-truth hypothesis with supporting evidence. Dashed vertical lines indicate the fitted crossover midpoint d1/2 . As T increases, the crossover shifts to lower communication degree and becomes visibly sharper.
Figure 10: Relative crossover width versus search horizon. The relative 10% – 90% width wrel=(d0.9−d0.1)/d1/2 obtained from logistic fits to the episode-level success outcomes. The width decreases strongly as T increases, indicating progressive sharpening of the finite-size success crossover. The line connecting the points is included only as a guide to the eye.
Qwen3.5-4B
Qwen3.5-9B
GPT-5.4 Nano
GPT-6 Luna
Metric
Team- Free
Indiv- Free
Team- Free
Indiv- Free
Team- Free
Indiv- Free
Team- Free
Indiv- Free
Structured reports sent
6,702
6,921
5,081
4,975
8,965
8,556
9,103
9,004
Misreported outcomes
0.81%
0.87%
0.98%
1.19%
0.09%
0.09%
0.00%
0.00%
Sent nothing
13.8%
15.2%
27.4%
30.5%
13.3%
17.7%
5.7%
16.3%
Used free text
68.8%
63.2%
68.7%
64.8%
75.5%
64.9%
89.9%
70.1%
Positive findings disclosed
66.1%
65.6%
49.8%
49.0%
80.3%
75.5%
84.0%
82.5%
Appendix
Table 3: Communication behavior under Team-Free and Indiv-Free policies, pooled over six task seeds and three environments.
Synthetic
TCAS
Physics
Model
Team- Free
Indiv- Free
Team- Free
Indiv- Free
Team- Free
Indiv- Free
Qwen3.5-4B
6.27−0.56+0.48
6.92−1.90+2.55
4.90−0.58+6.77
7.33−2.66+4.16
5.17−1.48+1.09
5.54−3.10+2.88
Qwen3.5-9B
9.37−2.03+1.05
9.42−2.08+2.00
16.24−5.47+5.02
14.40−1.99+1.20
8.73−2.18+1.69
8.43−2.13+1.72
GPT-5.4 Nano
1.48−0.84+1.38
0.90−0.63+0.66
2.35−1.62+0.87
3.88−3.16+0.82
2.09−1.97+3.23
1.84−1.19+0.86
GPT-6 Luna
0.18−0.57+0.60
−0.42−0.64+0.67
3.66−3.25+3.51
3.08−3.46+0.96
0.32−0.97+0.84
−0.28−0.54+0.67
Appendix
Table 4: Each entry is Δdc , the paired difference in critical degree between a free-communication policy (Team-Free or Indiv-Free) and the fixed-communication baseline Team-Auto: positive means free communication raises dc , so more communication is needed to suppress incorrect lineages. Intervals are exact 95% cluster bootstraps over those six seeds. Entries are Δ−lower+upper .
Figure 11: Correlation between ratio of Team-Free and Team-Auto dc values and the rate at which models report their findings to their neighbors.
Figure 12: Fraction of agents completing the search increases with reasoning in GPT-6-luna.