Discovered, Not Designed: Population Evolution for Collaborative and Compute-Intensive Model Discovery
Organizations: Meta
Abstract
LLM-driven evolution enables iterative model development, but two practical goals remain underexplored: finding model designs that transfer across related tasks and sustaining improvement when training is expensive. We introduce Population Evolution (PE), a collaborative, hierarchical framework that connects ongoing local searches through shared experimental evidence. PE evaluates code changes across related training instances and shares the results to guide subsequent proposals and promotion to larger training scales. For expensive targets, PE searches small training subsets and screens candidates through peer and intermediate evaluations before full-target training. We introduce RMD-Bench to evaluate both settings across ranking, watch-time prediction, RL algorithm discovery, and LLM/VLM pretraining. Compared with standalone evolution at matched source iterations, PE raises mean best local gains from 7.01% to 8.97% in ranking and from 2.84% to 3.85% in watch-time, while improving the best larger-scale outcome in all three joint-discovery families. In watch-time discovery, PE improves best larger-scale gains with four of five harnesses and all four proposers. On new recommendation datasets under shared target-side calibration, every evaluated PE design improves over the reference in mean performance. Under matched total GPU compute, completed LLM discovery runs yield a best relative accuracy gain of 2.48% and 13 successful candidates for PE, versus 0.92% and none for direct evolution. VLM loss reduction reaches 8.78% versus 5.05% under matched total GPU compute.
Figures & tables
| Evaluation support | MLE- bench ( 2024 ) | Innovator Bench ( 2025 ) | Research Gym ( 2026 ) | PostTrain Bench ( 2026 ) | MLS- Bench ( 2026 ) | Edge Bench ( 2026 ) | RMD- Bench (ours) |
| Collaboration protocol | — | — | — | — | — | — | |
| Cross-dataset transfer | — | — | — | — | |||
| Cross-scale transfer | — | — | — | — | — | ||
| Hierarchical training | — | — | — | — | — | ||
| Discovery yield vs. compute | — | — | — | — | — | — | |
| Calibrated starting configurations | — | — | — | — |
| Dataset | S | PE |
|---|---|---|
| KuaiLive | 8.93 | 23.19 |
| KuaiRand-1K | 5.53 | 9.29 |
| KuaiRand-Pure | 8.03 | 6.07 |
| KuaiRec | 1.03 | 2.42 |
| KuaiSAR | 11.52 | 3.88 |
| Mean | 7.01 | 8.97 |
| Target metric | Candidate-set mean | Best candidate | ||
|---|---|---|---|---|
| S | PE | S | PE | |
| KuaiVideo AUC (%) | 74.39 | 74.45 | 74.43 | 74.46 |
| VK-LSVD MAE (s) | 14.55 | 14.54 | 14.46 | 14.53 |
| Domain | Method | Full-target attempts | Successful candidates | Best gain (%) |
|---|---|---|---|---|
| LLM | Direct | 70 | 0 | 0.92 |
| PE | 39 | 13 | 2.48 | |
| VLM | Direct | 100 | 12 | 5.05 |
| PE | 61 | 57 | 8.78 |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Task family | Starting method | Discovery space | Source instances | Final-stage evaluation |
| Track A: joint discovery across related instances | ||||
| Ranking | FuXi-Linear ( Ye et al., 2026 ) | Sequence representations; temporal and retention mechanisms | KuaiLive, KuaiRand-1K/Pure, KuaiRec, KuaiSAR | KuaiRand-27K; 146.82M interactions; ranking score |
| Watch-time | EGMN ( Zhao et al., 2025 ) | Distributional head; duration prediction | CIKM16, KuaiRec, WeChat, KuaiRand-Pure/1K | KuaiRand-27K; 23.99M examples; MAE (s) |
| RL algorithm | Qwen2.5-Coder ( Hui et al., 2024 ) ; GRPO | Policy loss; rollout filtering | Five sources, different seeds; 1.5B policy, 15 steps | 3B policy, 50 steps; 423 LiveCodeBench problems; pass@1 |
| Track B: decomposed discovery for expensive targets | ||||
| LLM pretraining | Gemma 3-style ( Gemma Team, 2025 ) ; 185M | Architecture; optimizer; schedule; training mechanisms | FineWeb-Edu 100B core set; two 2% and three 5% views | Full core set; seven-task accuracy |
| Dataset | Users | Interactions | Items | Mean seq. length | Context length | Batch size per GPU | Epochs |
|---|---|---|---|---|---|---|---|
| KuaiLive | 22,497 | 4,381,721 | 151,849 | 194.77 | 512 | 47 | 30 |
| KuaiRand-1K | 988 | 290,969 | 2,025 | 294.50 | 1,024 | 38 | 30 |
| KuaiRand-Pure | 25,220 | 1,235,423 | 6,073 | 48.99 | 512 | 88 | 30 |
| KuaiRec | 7,176 | 12,529,113 | 9,958 | 1,745.97 | 1,024 | 40 | 50 |
| KuaiSAR | 17,491 | 1,595,760 | 55,505 | 91.23 | 512 | 46 | 50 |
| KuaiRand-27K | 27,285 | 146,821,616 | 461,994 | 5,381.04 | 1,024 | 51 | 50 |
| Component | Setting |
|---|---|
| Encoder | Four blocks; four attention heads; |
| Embeddings and position | Item dimension 128; positional linear channel dimension 32 |
| Temporal channel | Eight heads; base 2, stride 3, start index 10; learnable decay |
| Sequence computation | Chunk size 128; RoPE disabled |
| Sampled-softmax loss | 128 negatives; temperature 0.05 |
| Execution | Eight-GPU DDP; BF16/TF32 |
| Dataset | Train | Eval. | Total | User/query vocabulary | Item vocabulary | p99 scale (s) |
|---|---|---|---|---|---|---|
| CIKM16 | 246,853 | 61,714 | 308,567 | 9,837 | 122,642 | 5.237 |
| KuaiRec | 980,676 | 245,169 | 1,225,845 | 7,176 | 9,599 | 47.697 |
| 5,059,605 | 1,264,902 | 6,324,507 | 19,994 | 95,511 | 165.364 | |
| KuaiRand-Pure | 1,871,608 | 467,902 | 2,339,510 | 27,279 | 7,268 | 182.240 |
| KuaiRand-1K | 6,951,281 | 1,737,821 | 8,689,102 | 1,000 | 2,926,199 | 188.515 |
| KuaiRand-27K | 19,193,265 | 4,798,317 | 23,991,582 | 27,285 | 5,489,023 | 187.368 |
| Dataset | Learning rate | Weight decay | Dropout | ||
|---|---|---|---|---|---|
| CIKM16 | 0.091684 | 0.659274 | 2.023865 | 0.141097 | |
| KuaiRec | 0.25076 | 0.619568 | 3.41108 | 0.059406 | |
| 0.085141 | 0.454659 | 3.263077 | 0.055439 | ||
| KuaiRand-Pure | 0.064052 | 0.521568 | 4.296266 | 0.077754 | |
| KuaiRand-1K | 0.134232 | 0.646871 | 1.511465 | 0.362471 | |
| KuaiRand-27K | 0.102368 | 0.489522 | 1.778713 | 0.138667 |
| Property | Source (each of five instances) | Larger scale |
|---|---|---|
| Policy parameters | 1.5B | 3B |
| GRPO steps | 15 | 50 |
| Prompts per step | 64 | 32 |
| Evaluation problems | 408 | 423 |
| Evaluation tests | Public tests | Private tests |
| Training seed | Instance-specific | Fixed |
| Dataset | Role | Shards | Documents | Tokens | Steps |
|---|---|---|---|---|---|
| D2-A | Source | 36 | 1,924,096 | 1,925,811,698 | 694 |
| D2-B | Source | 36 | 1,915,904 | 1,921,710,324 | 694 |
| D5-A | Source | 91 | 4,853,760 | 4,862,315,015 | 1,750 |
| D5-B | Source | 91 | 4,844,544 | 4,859,363,728 | 1,750 |
| D5-C | Source | 91 | 4,853,760 | 4,863,973,228 | 1,750 |
| D10 clean | Intermediate | 182 | 9,712,640 | 9,727,423,143 | 3,470 |
| Dataset | Role | Shards | Shard indices | Raw rows | Steps |
|---|---|---|---|---|---|
| D2-A | Source | 20 | 0–19 | 48,500 | 48 |
| D2-B | Source | 20 | 50–69 | 48,500 | 48 |
| D5-A | Source | 50 | 0–49 | 121,250 | 118 |
| D5-B | Source | 50 | 50–99 | 121,250 | 118 |
| D5-C | Source | 50 | 100–149 | 121,250 | 118 |
| D10 clean | Intermediate | 100 | 150–249 | 242,500 | 237 |
| Domain | Scientific and computational constraints |
|---|---|
| Ranking | Avoid attention whose cost is quadratic in the full sequence length. |
| Watch-time | Preserve non-negative distributional support and stable likelihoods; treat CIKM16 duration as session context. |
| LLM | Distinguish vector and matrix parameters in Muon and maintain stable optimizer updates. |
| VLM | Preserve causal attention and the fixed visual interface. |
| RL | Preserve response groups and masks, and compute stable probability ratios. |
| Family | Instances | Trial cap |
|---|---|---|
| Ranking | Five sources and KuaiRand-27K | 50 |
| Watch-time | Five sources and KuaiRand-27K | 50 |
| LLM pretraining | Two 2% sources and two intermediate instances | 20 |
| LLM pretraining | Three 5% sources | 10 |
| VLM pretraining | Five sources and two intermediate instances | 20 |
| Family / tier | Parameters and ranges |
|---|---|
| Ranking | Dropout ; learning rate (log); per-GPU batch size (log integer); weight decay (log). |
| Watch-time | Learning rate (log); weight decay (log); ; ; dropout . |
| LLM: source, intermediate | Peak learning rate (log); weight decay (log); warm-up steps (log integer); decay fraction ; Muon momentum ; Q/K clipping threshold (log). |
| VLM: source, intermediate | Same ranges as LLM except warm-up steps: for 2% sources, for 5% sources, and for 10% intermediate instances (log integer). |
| Dataset | Dropout | Learning rate ( ) | Batch size | Weight decay ( ) |
|---|---|---|---|---|
| KuaiLive | 0.49 | 2.5 | 47 | 4.8 |
| KuaiRand-1K | 0.352 | 3 | 38 | 1.1 |
| KuaiRand-Pure | 0.28727 | 7.969455 | 88 | 6.251374 |
| KuaiRec | 0.3 | 4.3 | 40 | 14 |
| KuaiSAR | 0.53 | 5.2 | 46 | 9 |
| KuaiRand-27K | 0.378871 | 3.548962 | 51 | 1.985122 |
| Instance | ( ) | |||||
| LLM pretraining | ||||||
| D2-A | 4.480393 | 0.018409 | 66 | 0.143635 | 0.985564 | 114.656 |
| D2-B | 5.4 | 0.05 | 170 | 0.085 | 0.919 | 170 |
| D5-A | 4.72 | 0.022 | 78 | 0.108 | 0.982 | 121 |
| D5-B | 4.480393 | 0.018409 | 66 | 0.143635 | 0.985564 | 114.656 |
| D5-C | 4.480393 | 0.018409 | 66 | 0.143635 | 0.985564 | 114.656 |
| Family / sources | Enabled | Extra trials | Paired control | Total evaluations |
|---|---|---|---|---|
| Ranking / KuaiRand-1K | Yes | 3 | Up to 1 | Up to 5 |
| Ranking / other four sources | No | 0 | 0 | 1 |
| Watch-time | No | 0 | 0 | 1 |
| RL algorithm | No | 0 | 0 | 1 |
| LLM pretraining | No | 0 | 0 | 1 |
| VLM pretraining | No | 0 | 0 | 1 |
| Task family | GPU-hours per complete discovery run |
|---|---|
| Ranking | Hundreds |
| Watch-time prediction | Tens |
| RL algorithm discovery | Hundreds |
| LLM pretraining | Thousands |
| VLM pretraining | Thousands |
| Family | Source publication | Peer evaluation | Larger scale / intermediate | Full target |
|---|---|---|---|---|
| Ranking | 2% relative | 1% relative | 2% relative | — |
| Watch-time | 0.5% relative | 0.5% relative | 0.5% relative | — |
| RL algorithm | 10% rel.; abs. | 5% rel.; abs. | 4% rel.; abs. | — |
| LLM | 2% relative | 1% relative | 2% on both intermediates | 1% relative |
| VLM | 10% relative | 10% rel.; 0.02 abs. | 6% rel.; 0.02 abs., both intermediates | 4% rel.; 0.015 abs. |
| Domain | Dataset / summary | S | PE | PE best including peer candidates † |
|---|---|---|---|---|
| Ranking | KuaiLive | +8.93% | +23.19% | +23.19% |
| KuaiRand-1K | +5.53% | +9.29% | +10.66% | |
| KuaiRand-Pure | +8.03% | +6.07% | +6.07% | |
| KuaiRec | +1.03% | +2.42% | +6.47% | |
| KuaiSAR | +11.52% | +3.88% | +10.06% | |
| Mean | +7.01% | +8.97% | +11.29% |
| Domain | Dataset | Local PE minus S | PE including peers minus S |
|---|---|---|---|
| Ranking | KuaiLive | +14.26 | +14.26 |
| KuaiRand-1K | +3.76 | +5.13 | |
| KuaiRand-Pure | |||
| KuaiRec | +1.39 | +5.44 | |
| KuaiSAR | |||
| Watch-time | CIKM16 | +5.23 | +5.23 |
| Family | Source / iteration | Principal intervention | Final-stage outcome |
|---|---|---|---|
| Ranking | KuaiRand-Pure / 57 | Gated length-64 causal attention alongside linear retention; inherited temporal and spectral components. | KuaiRand-27K: 3.02% ranking gain. |
| Watch-time | WeChat / 21 | Duration-conditioned, left-truncated Laplace mixture with a learned near-median quantile readout. | KuaiRand-27K: 2.17% MAE reduction. |
| RL algorithm | Source 2 / 87 | Bounded advantage weights based on trajectory distance from the reference policy. | 3B, 50 steps: 6.64% pass@1 gain. |
| LLM | D5-B / 89 | EMA/SWA checkpoint averaging over approximately the final 20% of training. | Full target: 2.48% accuracy gain. |
| VLM | D5-A / 86 | Gated global image-context injection into patch tokens, combined with square-root warm-up. | Full target: 7.51% loss reduction. |
| Design | Metric | Parent score | Candidate score | Gain (%) |
|---|---|---|---|---|
| LLM, D5-B / 89 | Full-target accuracy (%) | 41.13 | 41.25 | 0.28 |
| VLM, D5-A / 86 | Full-target loss | 1.1200 | 1.1164 | 0.33 |
| RL, source 2 / 87 | Source pass@1 (%) | 3.43 | 3.19 |
| Best gain on each source dataset | Larger scale KuaiRand-27K | |||||||
| Configuration ( ) | Method | CIKM16 | KuaiRec | KuaiRand Pure | KuaiRand 1K | Best gain | Passed / tested | |
| A. Search harnesses (GPT-5.5; high) | ||||||||
| AdaEvolve (2) | S | 7.32 | 0.40 | 0.80 | 0.75 | 4.93 | 1.77 | 4/10 |
| PE | 12.55 | 0.67 | 0.81 | 0.16 | 5.08 | 2.17 | 5/7 | |
| EvoX (4) | S | 7.01 | 0.28 | 0.49 | 0.30 | 1.87 | 1.45 | 4/20 |
| PE | 11.01 | 0.58 | 0.49 | 0.11 | 3.01 | 1.05 | 18/18 | |
| Relative gain threshold (%) | PE | Direct |
| 0.50 | 21 | 6 |
| 1.00 | 13 | 0 |
| 1.50 | 8 | 0 |
| 2.00 | 3 | 0 |
| Component | Specification |
|---|---|
| Examples (train / val. / test) | 7,505,820 / 1,516,831 / 1,275,643 |
| Chronology | 5/1/1-week split after historical-feature weeks 0–19; training uses weeks 20–24 |
| Numerical features | Video duration and 31 other features; normalization constant: 157 seconds |
| Categorical features | Age, gender, geographic region, place, platform, and user agent |
| Duration encoding | 50 quantile-based buckets fitted on training data |
| Shared EGMN encoder | 16-dimensional categorical embeddings, scalar linear projections, and a 256/128/64 MLP |
| Method | Candidate / source | MAE (s) | Paired reduction (%) |
|---|---|---|---|
| Reference | EGMN | — | |
| PE | e33792c7 † | ||
| PE | 2f3242e3 † | ||
| PE | 92f5eb20 † | ||
| PE | e0b2ffa7 | ||
| PE | 94db2d04 |
| Component | Specification |
|---|---|
| Candidate rows (train / test) | 10,931,092 / 2,730,291 |
| Requests (train / test) | 803,007 / 209,728; 10,000 users per split |
| Item vocabulary | Minimum training count 10; 632,404 known items, plus padding and unknown entries |
| Visual features | 64-dimensional vectors for 3,242,315 items; frozen, with separate padding and unknown entries |
| Item representation | Learned item-ID embedding plus a trainable, bias-free projection of its visual vector |
| History encoding | Shared sequence encoder and embeddings for positive and negative histories; at most 100 items each |
| Parameter | Search range | Shared selected value |
| VK-LSVD | ||
| Learning rate | 0.076611 | |
| Weight decay | ||
| Entropy coefficient | 0.37454 | |
| Regression coefficient | 4.7585 | |
| Dropout | 0.365997 | |
| Model | AUC (%) | User gAUC (%) | Log loss |
|---|---|---|---|
| Reference | |||
| Best S | |||
| Best PE |
| Candidate | AUC (%) | User gAUC (%) | Log loss |
|---|---|---|---|
| Reference | |||
| b3cb61d8 | |||
| 5f007905 | |||
| 7ac9ec81 | |||
| e7ec2459 | |||
| 2755e3d1 |
| Source | Candidate | AUC (%) | User gAUC (%) | Log loss |
|---|---|---|---|---|
| — | Reference | |||
| KuaiRand-1K | 560d182a | |||
| KuaiRand-1K | 35335eb3 | |||
| KuaiRand-1K | e5c01c96 | |||
| KuaiRand-1K | 5d08ec24 | |||
| KuaiSAR | 3a70637c |
| Method | Candidate | Solved / 632 in each rerun | Pass@1 (%) |
| Reference | — | 145, 150, 146 | |
| PE | 93d95a15 | 146, 143, 139 | |
| PE | cbf1c653 | 146, 144, 148 | |
| S | a0936bce | 141, 143, 142 | |
| S | db21a301 | 146, 145, 144 | |
| S | 77431a4e | 136, 140, 136 |