An evolutionary origin of collective decision making in humans and machines
Authors: Guocheng Wang, Qi Su, Joshua B. Plotkin
Organizations: Department of Biology, University of Pennsylvania, Philadelphia, PA 19104, USA · School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai 200240, China · Meiji Institute for Advanced Study of Mathematical Sciences, Meiji University, Nakano, Tokyo 164-8525, Japan
Groups of individuals can solve collective problems more accurately than any single member, by aggregating their opinions. Recent theoretical work has identified individual-level reward schemes that allow uninformed individuals to evolve collective intelligence from the bottom up, through social learning. Yet these results are restricted to linear prediction problems and simple averaging, while the decision tasks that real groups confront are often non-linear, and the institutions that aggregate opinions are seldom single-layer averages: districts elect representatives who in turn vote on policy, referees advise editors who decide on publication. Here we develop a framework for the evolution of collective intelligence in multi-layer voting populations, where individuals observe limited information and groups recursively aggregate their opinions by majority rule. We prove that single-layer voting cannot solve non-linear classification problems under any individual reward scheme. We then identify a "marginal feedback" payoff structure, which rewards individuals only when their opinion is pivotal in their group, and at every layer above them. This reward scheme induces a layered population to evolve accurate collective solutions to complex, non-linear decision tasks through individual-level peer imitation alone. The collective behavior that emerges is equivalent to a multi-layer perceptron in machine learning. Our results provide a naturalistic account of hierarchical institutions, in which the outsize importance of swing voters is the incentive that sustains collective accuracy; and they identify the credit-assignment rule in machine learning as not just an engineered solution but a natural evolutionary outcome.
Figures & tables
Figure 1: A model of collective prediction by layered voting. The population predicts a binary outcome by aggregating the opinions of individuals, either in a single vote or in layers. ( A ) A binary outcome Y is determined by m factors through a fixed but unknown true classification rule Y=sign{f(X1,…,Xm)} . The factors are redrawn at random in each round, so the outcome varies over time. ( B ) Each individual observes a single factor XeA and issues an opinion YA=cAxeA , encoded by their strategy (eA,cA) : which factor to observe, eA , and the believed correlation between that factor and the outcome, cA . ( C ) Under single-layer voting the collective prediction is the sign of the total opinion, Y^=sign{∑AYA} . ( D ) Under two-layer voting each group of individuals first returns a majority verdict, and then the majority of the group verdicts produces the collective prediction.
Figure 2: A two-layer population whose members are rewarded only for pivotal opinions evolves, through peer imitation alone, to solve non-linear prediction problems. Panels A to D show the collective accuracy R as individuals in the population evolve their strategies in response to marginal feedback rewards; panels E to L compare a prediction sampled from an evolved two-layer population to the true outcome, with predictions shown in blue ( +1 ) and white ( −1 ). ( A and B ) Single-layer voting under the feedback reward scheme evolves near-perfect accuracy for a linear prediction problem ( A ) but performs no better than chance in a non-linear XOR problem ( B ). The linear problem is shown for m=10,50,100 factors and the non-linear XOR problem for m=2,3,4 . ( C and D ) Two-layer voting with the marginal feedback reward scheme evolves accurate prediction for both linear ( C ) and non-linear ( D ) problems. Accuracy on the non-linear problem declines as m grows, because m sets both the number of factors and the degree of the classification boundary. ( E to L ) The stationary prediction of a two-layer population ( F , H , J , and L ) closely matches the true outcome ( E , G , I , and K ) as seen in four two-factor problems: a linear boundary Y=sign{x1+x2} ( E and F ), a product boundary Y=sign{(x1−0.2)(x2+0.4)} ( G and H ), a parabola Y=sign{x12−2x2−3} ( I and J ), and a circle Y=sign{x12+x22−6} ( K and L ), reaching accuracies of 92 – 97% . Parameters: N≈104 (so that ngZ is close to 104 ), ng=101 ( C to L ), u=0.02 , ai∼U(−1,1) (uniform distribution), xi∼U(−3,3) , ci∼U(−3,3) , s=103 .
Figure 3: An intermediate number of groups predicts best, by balancing two kinds of diversity. For a fixed population size ( N=104 ) we vary the number of groups ng and measure the stationary accuracy ( A and B ) and the diversity of predictions ( C ). ( A ) For a linear prediction problem, accuracy is highest for a single group and declines only slightly as the group number increases. ( B ) For the non-linear (XOR) problem, accuracy is maximized at an intermediate number of groups. ( C ) For the non-linear prediction problem with m=4 , within-group diversity (the entropy of individual opinions) falls while inter-group diversity (the entropy of group verdicts) rises, as the number of groups increases; both types of diversity are large only for an intermediate number of groups. Diversity is shown in the initial and in the stationary state. Parameters: N≈104 , ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) , s=103 , u=0.02 .
Figure 4: Deeper hierarchies solve more complex problems under global marginal feedback. ( A ) A three-layer population, illustrated with first-layer groups of size n1=5 , second-layer groups of size n2=5 , and a third-layer group of size n3=7 . The population size is N=n1n2n3 . Individual A lies in first-layer group x (opinion sum ϕx ), and x lies in second-layer group y (verdict sum ϕy ). Under global marginal feedback A is rewarded only when their opinion is pivotal at both layers, that is when 1∣ϕx∣≤∣YA∣=1 and 1∣ϕy∣≤1=1 . ( B ) On the XOR problem ( m=4 ), evolution under global marginal feedback (Eq. 6 ) produces the highest accuracy; the two forms of local marginal feedback, each retaining a single pivotality indicator improve accuracy but less so; and feedback alone performs no better than random chance. ( C ) Stationary accuracy against the first-layer group size n1 , for three-layer populations with n3=5 , 11 and 25 (fixed N ). An intermediate group size is best, and most three-layer structures exceed the maximum accuracy of any two-layer structure in the same population (red dashed line). ( D ) Across problem complexities m , the best three-layer population ( n2=51 and n3=5 ) predicts more accurately than the best two-layer population ( ng=100 ). ( E ) Searching over all grouping structures, the maximum achievable accuracy grows with the number of layers: 0.57 for two layers, 0.87 for three, and 0.90 for four ( m=6 ). Parameters: N≈104 , m=4 ( B and C ), m=6 ( E ), u=0.02 , ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) , s=2×103 , n3=5 and n2=51 ( B and D ).
Figure S1: Collective intelligence is robust to the way individuals form opinions. In the main text an individual forms the continuous opinion YA=cAxeA ; here we vary how individuals form opinions and re-measure the stationary accuracy of two-layer voting, across different groupings. Faint lines reproduce the accuracy of the baseline model from panel A for comparison. ( A ) The baseline model, YA=cAxeA (the same data as Fig. 3 B). ( B ) Individuals give a binary opinion from the sign of the observed factor, YA=sign{cAxeA} ; discarding the magnitude of the factor lowers the stationary accuracy. ( C ) Individuals give a thresholded binary opinion, YA=sign{cA,1xeA+cA,2} , with strategy (eA,cA,1,cA,2) ; accuracy exceeds the baseline, especially for few groups. ( D ) Individuals give the sign of a quadratic function of the observed factor, YA=sign{cA,1xeA2+cA,2xeA} . ( E ) A population mixing all four ways of forming opinions, where imitation copies a neighbor’s method together with its parameters. Across all cases, two-layer voting evolves collective intelligence, with peak accuracy at an intermediate number of groups. The qualitative results are driven by the population and reward schemes, not by the specific way individuals form opinions. Parameters: N=104 (two layers), ai∼U(−1,1) , u=0.02 ( B ), xi∼U(−3,3) , ci∼U(−3,3) , s=103 .
Figure S2: Observing more factors raises the accuracy of collective prediction. Here each individual observes k factors and forms a linear opinion from them, YA=cA,1xeA,1+⋯+cA,kxeA,k , with strategy (eA,1,cA,1,…,eA,k,cA,k) . ( A to C ) Each individual observes one (the baseline model), two, and three factors, respectively. The more factors an individual observes, the more accurate the collective prediction, since individuals then process more information about the prediction problem. Parameters: N=104 (two layers), ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) , s=103 .
Figure S3: Collective accuracy persists under unequal group sizes. For a fixed population size N=10000 , group i is given size proportional to iα , so that α≥0 controls the inequality in group sizes and α=0 recovers equal groups. Two updating rules are shown: ( A ) an individual A is drawn from the population and then a partner B from A ’s group; ( B ) a group is drawn and then two individuals within it. The two updating rules produce qualitatively similar stationary accuracies. When groups are few, greater inequality lowers accuracy; when groups are many, greater inequality raises it, because concentrating individuals into fewer, larger groups restores the within-group diversity that small groups lack. Parameters: N=10000 , ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) , s=102 , u=0.02 , m=4 .
Figure S4: Both layered voting and marginal feedback are required to solve non-linear decision problems. Figure 2 shows that single-layer voting evolved under feedback fails on non-linear prediction problems while two-layer voting evolved under marginal feedback succeeds. Here we test the two remaining combinations: single-layer voting under marginal feedback ( A and B ) and two-layer voting under feedback ( C and D ). ( A and C ) For a linear prediction problem, both combinations evolve accurate prediction. ( B and D ) In a non-linear prediction problem, both fail to exceed random chance. Layered voting and marginal feedback are therefore each necessary, and neither alone is sufficient, for non-linear prediction problems. Parameters: N=104 , ng=101 , u=0.02 , ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) , s=103 .
Figure S5: Marginal feedback rewards only pivotal opinions. ( A ) Under the feedback payoff scheme, πA=YA(Y−Y^) , every individual is rewarded or penalized regardless of whether their opinion is pivotal. ( B ) The marginal feedback scheme, πA=YA(Y−Y^)1∣ϕ∣≤∣YA∣ , adds the indicator 1∣ϕ∣≤∣YA∣=1ϕ(ϕ−YA)≤0+1ϕ(ϕ+YA)≤0 , where the two terms cannot hold at once (except when ϕ=0 ). The first term marks an individual whose removal would flip the group verdict (individual A ); the second marks an individual a copy of whose opinion would flip it (individual B ). Only such pivotal individuals are rewarded (or penalized). For illustration a single group is shown, with ϕ the sum of all opinions (i.e., ϕ=ϕg(A) ).
Figure S6: Either component of marginal feedback promotes collective intelligence. Marginal feedback decomposes into two terms: YA(Y−Y^)1ϕg(A)(ϕg(A)−YA)≤0 and YA(Y−Y^)1ϕg(A)(ϕg(A)+YA)≤0 . The first term represents rewards for individuals whose opinion currently determines the group verdict (i.e., the group verdict will change if YA is removed); and the second term represents rewards for individuals whose opinion has potential to correct a wrong verdict (or disrupt a currently correct verdict) if a second copy of this opinion were added into the group. Thus, the first term measures individuals’ present decisiveness and the second term measures their potential decisiveness. Either component alone raises accuracy to the same stationary level as the combination of the two; the combination converges faster. Parameters: N=104 , m=3 , ng=101 , s=20 , u=0.02 .
Figure S7: Larger populations reach higher accuracy in non-linear prediction problems. Stationary accuracy against population size, for a non-linear prediction problem with m=2,3,4 factors. ( A ) A single-layer population cannot predict a non-linear prediction problem, so enlarging it does not help. ( B and C ) A two-layer population can be enlarged either by adding groups at fixed group size ( B ) or by enlarging groups at fixed group number ( C ); in both cases a larger population reaches higher stationary accuracy. Parameters: ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) , u=0.02 , Z=25 ( B ), ng=100 ( C ), s=103 .
Figure S8: The optimal number of groups grows slowly with population size. ( A ) The number of groups that maximizes stationary accuracy increases with population size, and decreases with the number of factors at fixed population size. ( B ) The corresponding optimal group size grows only slowly with population size, or remains roughly constant. A growing population therefore improves its predictions mainly by forming new groups rather than by enlarging existing ones. Parameters: ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) , s=103 , u=0.02 .
Figure S9: Global marginal feedback in multi-layer populations. ( A ) An L -layer population: in layer K , every nK entities form a higher-layer group whose verdict passes to layer K+1 , with N=n1n2⋯nL . Writing gK(A) for the layer- K entity containing individual A , global marginal feedback is πA=YA(Y−Y^)1∣ϕg2(A)∣≤∣YA∣1∣ϕg3(A)∣≤1⋯1∣ϕgL(A)∣≤1 , with L−1 pivotality indicators. ( B and C ) The probability that a random individual receives a nonzero payoff, given YA(Y−Y^)=0 . In the two-layer case ( B ) it rises with the number of groups and is independent of the problem scale m ; in the three-layer case ( C ) it is set mainly by the number of top-layer entities (i.e., n3 ). Individuals in deeper hierarchies are rewarded less often, because of the extra indicator. ( D ) In two-layer populations, the stationary payoffs are small due to the pivotality indicator, especially for few groups. So stronger selection amplifies payoff differences and raises accuracy until saturation; with more groups, a lower selection intensity suffices. Main-text results are in the saturation regime. Parameters: N=104 , u=0.02 , m=2 ( C and D ), ai∼U(−1,1) , xi∼U(−3,3) , ci∼U(−3,3) .
Figure S10: Collective accuracy is robust to the innovation rate. Stationary accuracy against the innovation rate u , for two-layer populations with few groups ( ng=51 ), an intermediate number ( ng=401 ), and many groups ( ng=1001 ). Innovation adds strategic diversity within groups. When groups are few and large, within-group diversity is already high, so accuracy is nearly independent of the innovation rate, and a very high rate lowers it by adding noise. When groups are many and small, within-group diversity is low, so a higher innovation rate raises accuracy, until a very high rate again adds noise. Parameters: N=104 , m=4 , s=103 .
Figure S11: Two forms of marginal feedback reach the same accuracy. The marginal feedback of Eq. 5 and the global marginal feedback of Eq. 6 follow from requiring the Lyapunov function V=E[(Y^−Y)ϕ] to decrease; we call them V -induced. Requiring instead that U=E[(Y^−Y)2]=4(1−R) decrease yields U -induced forms, equal to the V -induced forms times an extra indicator 1∣ϕ∣≤1 on the top-layer margin (see supplementary text section S4). ( A and B ) The two forms reach the same stationary accuracy in a two-layer ( A ) and a three-layer ( B ) population, so the extra top-layer indicator is redundant; the V -induced form, with one fewer indicator, pays individuals more often and converges faster. Parameters: N=103 , m=3 ( A ), m=4 ( B ), ng=101 ( A ), n2=51 and n3=5 ( B ), s=103 .
Figure S12: The approximation of the sign function and its derivative. ( A ) Here we use function tanh(x/ϵ) to approximate the sign function. Parameter ϵ controls the characteristic length (bandwidth) of the transition region between −1 and 1 for tanh(x/ϵ) . ( B ) For the function tanh(x/ϵ) , we can compute its derivative. The derivative can be approximated by an indicator function 1∣x∣≤ϵ(x) , which equals 1 if ∣x∣≤ϵ and equals 0 otherwise.
Suppose a committee, expert panel, or other group is making judgments on some issues, where these may be not just yes/no-questions, such as whether a defendant is guilty, but also variables with many possible values, such as macroeconomic or meteorological variables or travel directions. Furthermore, there may be interconnections between different issues, as in the case of economic or climate variables. How can the group arrive at "intelligent" collective judgments, based on the group members' individual judgments? We investigate three challenges raised by this judgment-aggregation problem. First, reasonable methods of aggregation (such as defining the collective judgment for each issue as the average or median judgment) can produce inconsistent collective judgments. Secondly, many methods of aggregation are manipulable by strategic voting. Finally, not all methods of aggregation are conducive to tracking the truth on the issues in question. We prove new impossibility or possibility theorems on all three challenges, identifying what it takes to produce collective judgments in a consistent, non-manipulable, and truth-tracking manner and thereby to achieve collective intelligence through aggregation. Overall, the median method, though imperfect, performs reasonably well. We also note the relevance of our analysis for non-human group decisions.
Franz Dietrich, Christian List
We thank two anonymous referees and an editor for very helpful comments.
Across millennia, complex societies have faced the same coordination problem of how to organize collective action among cognitively bounded and informationally incomplete individuals. Different civilizations developed different political institutions to answer the same basic questions of who proposes, who reviews, who executes, and how errors are corrected. We argue that multi-agent systems built on large language models face the same challenge. Their central problem is not only individual intelligence, but collective organization. Historical institutions therefore provide a structured design space for multi-agent architectures, making key trade-offs between efficiency and error correction, centralization and distribution, and specialization and redundancy empirically testable. We translate seven historical political institutions, spanning four canonical governance patterns, into executable multi-agent architectures and evaluate them under identical conditions across three large language models and two benchmarks. We find that governance topology strongly shapes collective performance. Within a single model, the gap between the best and worst institution exceeds 57 percentage points, while the optimal architecture shifts systematically with model capability and task characteristics. These results suggest that collective intelligence will not advance through a single optimal organizational form, but through governance mechanisms that can be reselected and reconfigured as tasks and capabilities evolve. More broadly, this points to a transition from \textbf{self-evolving agents} to the \textbf{self-evolving multi-agent system}. The code is available on GitHub.
More capable agents do not necessarily form a more capable collective. A multi-agent system may jointly possess sufficient information yet fail because evidence is poorly routed, unreliable reports enter public belief, correlated claims masquerade as independent support, shared state becomes stale or strategically distorted, or useful evidence is exposed through an ineffective action interface. We ask when additional resources should improve the reasoner and when they should instead change the institutional structure through which the collective forms and acts on public information. Drawing on functional distinctions from research on group decision making and distributed cognition, we construct controlled artificial ecologies around four loci of collective failure: access and routing, admission and dependence, state maintenance and incentives, and representation and action. Across these ecologies, we separately vary model capability and institutional structure, pairing positive interventions with matched reasoning baselines and mechanism-breaking controls. The experiments reveal a consistent boundary: institutions help when they repair failures in how a collective constructs usable public state, but lose their advantage when their signals are uninformative or uncheckable, when stronger intelligence can perform the same transformation directly, or when the resulting state cannot support reliable action. Our results recast the choice between intelligence and institutions as a diagnosis of where collective reasoning fails.
Zhengye Han
New York University Department of Electrical and Computer Engineering Brooklyn, New York, United States