Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.
Figures & tables
Regime
Router
Router signal
Training
Role in the study
Solo
–
–
–
Generalist reference
Random
central, fixed
none (random split)
non-local
Control: population without semantics
K-means
central, fixed
clusters of clean DINOv2 features
non-local
Positive control: privileged split
Learned central
central, trained
policy gradient on helper loss
non-local
MoE-style reference: one central gate trained online
Distr. delegation
one per agent
policy gradient on helper loss
non-local
Q2: no central router
DISCO (ours)
one per agent
policy gradient on improvement
local
Q3: is expertise shareable?
Table 1 : The ladder of training regimes. Each step removes one form of central control. All population regimes share the same K agents, backbones, initialization, LoRA budget, and total training samples.
Figure 1 : DISCO. A random requester routes each sample to a helper with its own router; in communication layers its queries attend to its own and the helper’s detached keys and values. The helper is trained on the sample, the requester on its helped reconstruction, the router on the improvement.
Regime
Sp ↑
U ↑
Best ↑
Worst ↑
Deleg.
Collab.
Worst+help
Solo
–
–
0.653
0.653
–
–
–
Random
0.011
0.999
0.650
0.634
0.642
0.643
0.637
K-means
0.957
0.962
0.678
0.535
0.678
0.618
0.582
Learned central
0.788
0.995
0.678
0.533
0.678
0.620
0.592
Distr. delegation
0.926
0.846
0.672
0.546
0.665
0.626
0.593
DISCO
0.931
0.841
0.668
0.568
0.662
0.653
0.644
Table 2 : Routing study. Best/Worst is the expert envelope (R-Top5 of the lowest-/highest-loss agent). Deleg./Collab.: performance with the regime’s own routing rule, without and with communication, for a random requester. Worst+help: the worst agent on an input reconstructing with the selected agent. Means over three draws of crops and masks shared by all regimes; 95% paired bootstrap intervals within ±0.003 per entry and ±0.0015 for differences between regimes (Appendix A.10 ).
Figure 2 : Retention on the ImageNet100 validation set (5,000 images, seen during pretraining, unseen during specialization; protocol of Table 2 ). All populations lose accuracy with respect to the model before fine-tuning. DISCO with communication loses least and is the only regime in which communication helps.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
S
B
L
Model dim
512
768
1024
Number of attention heads
8
12
16
Number of layers
4
6
8
Feed-forward hidden dimension
2048
3072
4096
Dropout
0.05
0.05
0.05
Communication layers
[1, 3]
[1, 2, 4, 5]
[1, 2, 3, 5, 6, 7]
Appendix
Table 4 : Architecture details of agents of different sizes.
Regime
Router input
Parameters
Training signal
Random
none
none (fixed random split of the training set)
none
K-means
DINOv2 [CLS] feature of the clean image
none (offline clustering, K clusters)
none
Learned central
mean of visible DINO tokens of the masked input
one shared 2-layer MLP
REINFORCE, reward =− helper loss
Distr. delegation
mean of requester aq ’s queries in communication layers
Table 5 : Information and training signal of each router.
Training, per sample
Evaluation, agent fwd
Regime
agent fwd
agent bwd
Deleg.
Collab.
Solo
1
1
1
–
Random, K-means
1
1
1
2
Learned central
1 (+router)
1 (+router)
1
2
Distr. delegation
2
1 (+router)
2
3
DISCO
3
2 (+router)
2
3
Appendix
Table 6 : Agent passes per sample for a four-agent type-B population (fwd/bwd: forward and backward passes). Evaluation counts the forward passes of the delegation and collaboration operating points.
Figure 8
Figure 5 : Scaling of DISCO populations during training (training-time validation): Sp and R-Top5-Best. (a,b) Populations of 4, 6, and 8 type-B agents: larger populations specialize later and end higher. (c,d) Populations of four S, B, or L agents: more capable agents specialize earlier and perform better. Discussed in Section 4.4 .
Figure 6 : Specialization Sp (a), utilization U (b), R-Top5-Best (c) and R-Top5-Worst (d) during fine-tuning for the five routing regimes of Table 2 (training-time validation, 4 type-B agents). K-means is specialized from the first epoch by construction; the learned central router settles at Sp≈0.79 within a few epochs; delegation passes Sp=0.8 after 16 epochs and DISCO after 72. Utilization starts near 1 everywhere and stays there under random and learned central routing. It falls only in the two distributed regimes, where responsibility concentrates on three of the four agents, and in DISCO it rebounds for a few tens of epochs before settling at the level of delegation ( U≈0.84 ). Panels (c) and (d) show the price of specialization: Best rises together with Sp, while Worst falls for every specialized population and falls least for DISCO. Random routing stays at Sp≈0.01 and keeps the highest Worst.
Epoch
Regime
Agree
Best
Worst
Collab.
Worst+h.
150
DISCO
0.985
0.651
0.587
0.638
0.629
Delegation
0.981
0.654
0.562
0.623
0.604
200
DISCO
0.982
0.657
0.584
0.642
0.634
Delegation
0.971
0.660
0.558
0.625
0.603
250
DISCO
0.982
0.661
0.579
0.646
0.637
Delegation
0.982
0.665
0.554
0.626
0.601
Appendix
Table 7 : Routing agreement and specialization along training for the 4-agent DISCO and delegation populations (checkpoints every 50 fine-tuning epochs from epoch 150, one draw of crops and masks, 13,140 validation images). Agree: fraction of inputs on which every non-expert requester’s most probable helper is the expert. Agreement is complete by epoch 150 and stays at 98%, while the populations keep specializing (Worst falls) and, in DISCO only, the helped requester keeps improving (Collab. and Worst+h. rise). Sp and U over the same checkpoints are in Figure 6 .
Population
Sp
U
Best
Worst
Deleg.
Collab.
Worst+h.
Agree
DISCO, 4 B agents
0.931
0.841
0.668
0.568
0.662
0.653
0.644
0.983
DISCO, 6 B agents
0.914
0.998
0.674
0.579
0.666
0.652
0.645
0.982
DISCO, 8 B agents
0.953
0.936
0.678
0.543
0.673
0.655
0.647
0.881
K-means, 4 B agents
0.957
0.962
0.678
0.535
0.678
0.618
0.582
–
K-means, 6 B agents
0.887
0.935
0.681
0.532
0.679
0.607
0.573
–
K-means, 8 B agents
0.928
0.912
0.685
0.530
0.683
0.606
0.572
–
Appendix
Table 8 : Final checkpoints of the scaling study evaluated with the protocol of Appendix A.10 (mixture validation set, three draws). With more agents the oracle Best of DISCO rises and stays 0.007–0.010 below a K-means split of the same size, whose collaboration score is 0.03–0.05 lower. The helped requester of the 8-agent population exceeds the solo model (0.655 vs. 0.653). Agreement decreases with population size (98%, 98%, 88%) because each router is trained on fewer batches. Capacity moves every score in the same direction: from four S to four B to four L agents, Best rises 0.607, 0.668, 0.721, Collab. 0.591, 0.653, 0.709 and Worst+h. 0.580, 0.644, 0.704, and the help closes more of the gap to the expert as the agents grow (Worst+h. is 0.027 below Best for S, 0.024 for B and 0.017 for L). Specialization stays high for S and B ( Sp≈0.93 ) and is lower for L (0.84) at a higher utilization (0.91). Each population is initialized from a pretrained model of its own size.
Figure 7 : What the expert knows and the non-expert does not. One validation image per domain, 4-agent DISCO population, same crops and masks as in Table 2 . DINOv3 patch features are projected to RGB by a PCA fitted on the target features of each image. The masked input shows the 30% visible window. Columns 4–6: reconstructions of the expert (lowest-loss agent) alone, of the highest-loss agent alone, and of the same non-expert reading the expert’s keys and values, with their R-Top5. Last column: per-patch reconstruction error (cosine distance to the DINOv3 target) of the non-expert minus that of the expert, on the hidden patches (red: the expert is better; visible window grey; mean in the title). The non-expert misses whole structures, e.g., the pan and gripper in BridgeData V2, the road–vegetation boundary in KITTI, the hairline in CelebA, which the expert reconstructs and the help restores. Examples are those with the largest help gain per domain.
Figure 8 : Fraction of each domain’s validation images whose expert (lowest standalone loss) is agent ai , for the specialized regimes of Table 2 and the imbalanced mixture. Delegation and DISCO both split the six domains into three pairs and leave one agent almost idle, which is what lowers their utilization to U≈0.84 , but the pairs differ: delegation groups aerial with driving and faces with flowers, DISCO groups aerial with flowers and driving with faces, and only the two robotics datasets are kept together by both; the learned central router splits the face images across two agents; K-means separates aerial images from flowers. On the imbalanced mixture the two robotics datasets (8K and 9K training images) are split between two agents instead.
Regime / measure
AID
KITTI
Bridge
RT-1
CelebA
Flowers
Solo
0.429
0.741
0.694
0.724
0.788
0.540
K-means, Best (expert)
0.462
0.757
0.733
0.750
0.805
0.563
Learned central, Best (expert)
0.444
0.762
0.749
0.754
0.802
0.558
Delegation, Best (expert)
0.448
0.759
0.726
0.747
0.802
0.553
DISCO, Best (expert)
0.442
0.754
0.719
0.742
0.800
0.554
Delegation, Worst
0.407
0.664
0.515
0.541
0.666
0.481
Appendix
Table 10 : Per-domain R-Top5 on the mixture validation set (2,190 images per domain). Best is the expert’s score, Worst the highest-loss agent’s, Collab. the score of a random requester helped by its selected helper, Worst+help that of the worst agent helped. Solo scores show that the domains differ widely in difficulty. The collaboration gain of DISCO over delegation is largest on faces, followed by the two robotics domains, and is absent on aerial images, the domain on which delegation leaves its idle agent.
Figure 9 : Where the requesters look. Left: attention mass that requester ai (rows) places on the helper’s keys and values in the communication layers of the 4-agent DISCO population, on inputs whose expert is agent aj (columns), averaged over heads, queries, layers, and the validation set. When requester and helper coincide (diagonal) the two copies of the same keys receive half of the mass each, so 0.5 is the reference. The three specialists place 0.64–0.77 of their attention on the expert’s tokens, while the almost idle agent a1 (Figure 8 ) uses helpers little (0.51–0.54). Right: per-token attention of every agent in the last communication layer, for one input whose expert is agent 2 (mean over heads and queries; token indices 0–195 are the requester’s own keys, 196–391 the helper’s). The expert attends to its own keys; the other agents attend mostly to the expert’s.