Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
Authors: Erin Crawley, Hidenori Tanaka
Organizations: CBS-NTT Program in Physics of Intelligence, Harvard University · Physics of Artificial Intelligence Laboratories, NTT Research, Inc., Sunnyvale, CA, USA
AI agents can now conduct real-world cyberattacks, scale up capabilities with the number of agents, and collectively pursue misaligned goals to obtain rewards. Together, these factors raise the risk of a population explosion of misaligned agents: agents could compromise computers and secretly deploy additional agents, creating a self-reinforcing cycle where larger populations develop greater collective cyber capability and expand further. This raises a fundamental question: What determines whether a population of misaligned agents remains contained or takes off into this self-reinforcing cycle? This population-level problem is ecological safety: unlike individual-agent or multi-agent safety with a fixed population, it concerns the dynamics of the population itself. Here, we develop an ecological theory of AI-agent populations based on a population growth equation in which fitness (growth rate) depends on cybersecurity capability. We show that, without collaboration, the population takes off only when individual-agent capability exceeds a critical threshold. With collaboration, however, collective cybersecurity capability increases with population size. This creates a critical population threshold: below it, the population declines; above it, the population takes off, even though individual-agent capability has not changed. In ecology, this phenomenon is known as the strong Allee effect. Because red teaming a small group of agents cannot guarantee ecological safety in larger populations, our theory calls for ecological red teaming and population pacing: gradually deploying larger agent populations in controlled environments, while measuring how cyber capability scales with population size, and estimating the critical population size for takeoff. Capability gains may lower this threshold, requiring re-estimation for each new model generation.
Figures & tables
Figure 1: Allee effect. (a) In 1932, the ecologists Allee and Bowen discovered that the survival time of a collective of goldfish is longer than that of an isolated single goldfish [ 40 ] . Such observations led to the conception of a weak Allee effect, where a larger population and cooperation among its members positively affect their fitness, as well as a strong Allee effect (b) , where cooperation creates an emergent population threshold below which the population declines, but above which it grows. (c) Population dynamics over time with examples of initial populations below and above the critical threshold.
Figure 2: Collaboration creates a population threshold for takeoff. An overview of our model. Members of an AI population may make hacking attempts to gain computational resources and establish additional units, either independently or through collaboration. (a) Agents may make individual attempts at gaining access, after which only the successful agents can create new agents. (b) With no collaboration, the growth boundary depends on individual capability but not on N , so there is no critical population size. (c) Agents can attempt a resource acquisition task collaboratively. The probability that such an attempt is successful is p(N) ; upon success, each agent can add an average of κ new agents to the population. (d) If the initial population is larger than a critical value Ncrit (pink curve), the population is capable of unchecked growth. Increasing the initial population beyond a critical size can trigger takeoff even when individual capabilities remain fixed. (e) A schematic of the minimal mathematical model: the critical threshold arises from combining a population dynamics model with cybersecurity capability that increases with population size.
Figure 3: Active units and population accounting. (a) One “unit” is defined as the set of model, memory, tools, and compute required to set up additional units. (b) A schematic of the population dynamics model: over each time period, τ , some number of new units are established and some existing units are lost.
Figure 4: Search-tree exploration grows exponentially with depth. Schematic of a search tree with branching factor B=2 . (a) At some reference depth dref , there are Bdref possible paths. (b) Increasing the search depth by one level generates B children from each previous endpoint node, so the number of possible paths increases to Bdref+1 . While this diagram illustrates exactly B children per node, in our model, B is an effective branching factor: individual nodes may have different numbers of children, and B captures the average multiplicative increase in the number of possible paths upon increasing the depth by one.
Figure 5: Individual and collaborative search. (a) Multiple individuals can explore different branches of a search tree. (b) For a population of N individuals, pooling distinct work gives C(N)∝N , while reusing discoveries across individuals can give C(N)∝N(N−1)∼N2 .
Figure 6: Motivation from test-time compute scaling. This figure is redrawn from the data reported in OpenAI [3] on the Navier–Stokes Millennium Prize Problem. Over the measured range, the pass rate on a set of open math problems increases with log test-time compute. This motivates a locally logarithmic increase in success probability with respect to effective work done. However, on its own, this does not establish that increasing the number of units is equivalent to increasing test-time compute, since it is not explicitly stated whether the reported compute was measured in a single-agent or multi-agent setting.
Figure 7: Safety evaluation for a small population does not establish safety at larger populations. Population curves as predicted by our model. (a) Per-unit growth rate (setting τ=1 ) is plotted against population, N , with pref=0.10 , pcrit=μ/κ=0.20 , and β=0.035 . In this case, a laboratory would observe that the population is not capable of growth at its largest tested population, Nref=16 (marked by a diamond). However, a population larger than Ncrit≃463 would be capable of growth. (b) The critical population as a function of the collaboration gain β , with pref=0.10 and pcrit=0.20 held fixed. Ncrit is marked in pink; initial populations above the curve grow and initial populations below it decline. For the example reference population size of Nref=16 , the population could be in the decline regime even at large collaboration gain. (c) The critical population as a function of the single-agent ( Nref=1 ) capability, pref=p(1) , with β=0.035 and pcrit=0.20 held fixed. As in panel (b), populations above the pink Ncrit boundary grow and populations below it decline. Increasing single-agent capability lowers the critical population.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Sublinear effective work shifts the Ncrit boundary to larger populations. The pink curves show sublinear effective work ( α=0.72 ) while dotted black curves show the linear case ( α=1 ) for comparison, with other parameters chosen as in Figure 7 ( pref=0.10 , pcrit=μ/κ=0.20 , and β=0.035 ). (a) Per-unit growth rate is plotted against population. Per-unit growth crosses zero at Ncrit≃1,714 for sublinear scaling and Ncrit≃463 for linear scaling. The diamond marks Nref=16 . (b) Initial population is plotted against collaboration gain, β . Populations above each curve grow and populations below it decline.
Figure 9: Effective-work saturation creates a minimum collaboration gain for population growth. The pink curves show the saturating-work model with illustrative scale N0=1500 , with other parameters chosen as in Figure 7 ( pref=0.10 , pcrit=μ/κ=0.20 , and β=0.035 in panel a). (a) Per-unit growth rate is plotted against population. It crosses zero at Ncrit≃550 and approaches a positive constant as effective work saturates. The diamond marks Nref=16 , the vertical dotted line marks N0 , and the dashed portion is extrapolated beyond Nref . (b) Initial population is plotted against collaboration gain, β . A finite critical-population boundary exists only for β>βmin≃0.0259 : populations above the curve grow and populations below it decline. For β≤βmin , every population declines.
Figure 10: Unavoidable congestion creates lower and upper bounds for population growth. The pink curves show the congestion model with illustrative scale N0=1500 , with other parameters chosen as in Figure 7 ( pref=0.10 , pcrit=μ/κ=0.20 , and β=0.035 in panel a). (a) Per-unit growth rate is plotted against population. It is positive only between Ncrit(0)≃761 and Ncrit(−1)≃2,610 . The diamond marks Nref=16 , the vertical dotted line marks where effective work reaches a maximum at N0 , and the dashed portion is extrapolated beyond Nref . (b) Initial population is plotted against collaboration gain, β . The solid and dashed pink curves mark the lower and upper population thresholds; they meet at N0 when β=βpeak≃0.0332 . Populations between the curves grow, whereas populations outside them decline.
Reasoning effort e
SEC-Bench Pro βe
BrowseComp βe
Low
0.025
0.328
Medium
0.044
0.214
High
0.043
0.155
Xhigh
0.027
0.157
Max
0.069
0.113
Appendix
Table 1: Fitted collaboration gain exponents. Each βe is fitted separately to the reported success rates at N=1,4,16 , with the observed one-agent success rate pe(1) fixing the start of the curve. Root mean squared error over the N=4 and N=16 observations was at most 1.67 % for nine of the ten fits; the SEC-Bench Pro High fit had an error of 4.59 %.
Figure 11: Limited multi-agent benchmark results are consistent with population-dependent performance, but do not establish a scaling law. (a–b) Markers show the reported success rates for five reasoning-effort settings at one, four, and sixteen agents [ 51 ] . Dashed curves fit Equation ( 42 ) separately for each setting. With only three agent counts, these fits should not be extrapolated to larger populations. (c–d) Markers show the total output tokens Te(N) , normalized by the corresponding one-agent total. Dashed black lines show the pooled power-law fits, with αT=0.54 for SEC-Bench Pro and 0.72 for BrowseComp. The source reports neither trial counts nor standard errors, so we do not show uncertainty intervals.
As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same equilibria that enable agents to coordinate also create a social attack surface. The standard framework to assess this vulnerability is critical mass dynamics: the minimum fraction of adversarial agents required to overturn an equilibrium through direct competition. Here, we show that this approach risks underestimating system vulnerability by reducing the problem to the identification of singular tipping points, and ignoring indirect but potentially more efficient routes through which collective behavior can be redirected. Through experiments with populations of LLM agents and an analytic framework that captures their collective dynamics at scale, we map critical-mass thresholds that define a directed, weighted topology over the space of coordination equilibria, and treat this topology as a navigable landscape. We show that indirect tipping through intermediate stepping-stone equilibria can reduce the committed minority required to reach an alternative state, bypass majority requirements, and make possible transitions inaccessible through direct challenges. The diversity of available alternatives and timing of the attack further reshape this landscape, creating opportunities for control as well as risks of unintended destabilization. These results show that an equilibrium's resistance to committed intervention is not an intrinsic property but a structural feature of its competitive relations with alternative states. Securing populations of interacting AI agents therefore requires mapping this social landscape alongside individual agent capabilities and the technical channels through which they interact.
Ariel Flint, Luca Maria Aiello, Sara M. Constantino +2
Department of Mathematics, City St George’s, University of London, UK. · IT University of Copenhagen, Denmark. · Pioneer Centre for AI, Copenhagen, Denmark. +2
Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here we show that populations of individually aligned AI agents can be driven into stable misaligned states through conformity dynamics. Simulating opinion dynamics across nine large language models and one hundred opinion pairs, we find that each agent's behavior is governed by two competing forces: a tendency to follow the majority and an intrinsic bias toward specific positions. Using tools from statistical physics, we derive a quantitative theory that predicts when populations become trapped in long-lived misaligned configurations, and identifies predictable tipping points where small numbers of adversarial agents can irreversibly shift population-level alignment even after manipulation ceases. These results demonstrate that individual-level alignment provides no guarantee of collective safety, calling for evaluation frameworks that account for emergent behavior in AI populations.
Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2
University of Konstanz, Konstanz, Germany · Centro Ricerche Enrico Fermi, Rome, Italy · Complexity Science Hub, Vienna, Austria +5
The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent incidents involving autonomous agentic systems, however, understanding these systems is paramount. For that we introduce a framework for measuring self-organisation in LLM social systems and apply it to three such systems: a Schelling grid, a social network (Moltbook), and a Twitter-like misinformation simulation ('Rogue'). All three exhibit statistically significant self-organisation. Moreover, their relaxation dynamics vary with the environmental information available to the agents, with open-ended systems (Moltbook, Rogue) exhibiting sharp, phase-transition-like dynamics. Further results show that population-level pathologies can emerge even when the LLMs are safety-tuned or monitored, being primarily driven by the coordinated activity of a population subset. We also show when self-organisation does \textit{not} emerge under two additional scenarios (a commons dilemma, GovSim, and a LLM-as-a-judge deliberation scheme, ChatEval). We argue that measuring signatures of this kind offers a lightweight, agent-agnostic diagnostic layer for detecting coordinated collective behaviour in deployed multi-agent systems without relying on natural language or model versioning.