RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models
Organizations: Meta
Abstract
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.
Figures & tables
| Iter. | Intervention | Outcome | Finding and claim status |
|---|---|---|---|
| 0 | Published HSTU anchor | 0.2098 | Published comparison point |
| Stage I — compose known components ( SYNAPSE ); one objective bet | |||
| 1–2 | SSD-style input compression (IC) PRISM user conditioning | 0.2140 | ; measured historical endpoint |
| 3 | multi-token scoring head | 0.2161 | SYNAPSE , : two pooled facets (recent intent, long-term taste), one ANN call; measured historical endpoint |
| 4 | Preference optimization (DPO/IPO/SimPO) | DPO: vs. its 0.2191 reference | All three used target-derived negatives (pre-fix mining) and hurt as preference accuracy approached 1; reward over-optimization remains a hypothesis. A leak-free re-derivation mining only internal positions is neutral ( ). Rejected |
| Stage II — refine (diminishing returns) | |||
| Dimension | Reported | Reference | Relative |
|---|---|---|---|
| ML-20M BASE | 0.1948 | 0.1895 p | |
| ML-20M LARGE | 0.2192 | 0.2098 p | |
| ML-32M BASE | 0.1562 | 0.1488 i | |
| ML-32M LARGE | 0.1726 | 0.1660 i |
| Dataset | Vanilla | SYNAPSE | UDK |
|---|---|---|---|
| Foursquare-TKY | 0.0181 | 0.0200 ( ) | 0.0202 ( ) |
| Foursquare-NYC | 0.0156 | 0.0180 ( ) | 0.0195 ( ) |
| Gowalla | 0.0441 | 0.0460 ( ) | 0.0467 ( ) |
| Yelp | 0.0321 | 0.0359 ( ) | 0.0382 ( ) |
| Condition | HSTU EA | HSTU CDR | LitGPT EA | LitGPT CDR | Cost/task | Tool actions |
|---|---|---|---|---|---|---|
| CC, one pass | ||||||
| CC, extended budget | ||||||
| CC best-of- + blind selection | ||||||
| CC CC CC | ||||||
| CC Codex CC | ||||||
| Codex CC Codex |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Control representation | Node granularity | Primary objective/setting | Relation to this work |
|---|---|---|---|---|
| StateFlow ( Wu et al., 2024 ) | Programmed finite-state workflow | LLM instructions, tools, and functions | Task success/cost on interactive SQL, Bash, and ALFWorld | Establishes state-driven control and state-specific instructions; EOP does not claim either abstraction |
| GPTSwarm ( Zhuge et al., 2024 ) | Optimizable computational graph | Operations and recursively composed agent graphs | Prompt/connectivity optimization on reasoning tasks | Establishes agents-as-graphs and diversity precedent; motivates a stricter correctness/decorrelation test |
| ADAS / AFlow ( Hu et al., 2025 ; Zhang et al., 2025 ) | Agent code or code-represented workflow search | Generated agents, prompts, operators, and edges | Automatic agent/workflow discovery across benchmark tasks | RankEvolve uses a fixed human-authored flow and makes no automatic-design claim |
| LangGraph ( LangChain, 2024 ) | Stateful graph runtime | Arbitrary callables, subgraphs, or wrapped agents | Durable production orchestration | Provides overlapping runtime primitives; EOP is an ML-evolution authoring/adapter instantiation |
| RankEvolve (this work) | Semi-structured EOP compiled to runtime state | Black-box coding products at coding/review nodes; functions elsewhere | Long-horizon ML evolution and executable patch correctness | Claimed delta: product boundary, domain specialization, and deployment-seeded, budget-controlled evidence with a pre-specified transfer split |
| Parameter | BASE | LARGE |
|---|---|---|
| Blocks | 4 | 16 |
| Heads | 4 | 8 |
| 64 | 32 | |
| 256 | 256 | |
| Seq. length | 200 500 | 200 500 |
| Negatives | 128 | 128 |
| Construct | Compile/static check | Runtime behavior | Not guaranteed |
|---|---|---|---|
| Phase dependency | Referenced phases exist; DAG outside explicit loops | Successors remain unavailable until predecessor completion | That the predecessor’s artifact is semantically correct |
| Human gate | Gate has prompt and legal continuations | Transition blocks until a recorded approve/reject event | That the human decision is correct |
| Branch/join | Source variable, cardinality, and join policy are declared | Supported node backend instantiates one branch per item and records join state | Fan-out on an untested backend or useful diversity across branches |
| Loop/budget | Target phase exists; exit and hard bound are present | Attempt and spend counters stop further dispatch at the bound | That an agent chooses the best stopping point |
| Tool requirement | Tool name/schema resolve before launch | Calls outside the allowlist are rejected and recorded | Correct tool arguments or scientifically valid interpretation |
| Checkpoint/restart | Serializable state and artifact references | Resume begins at the last committed node boundary | Recovery inside an uncommitted third-party agent invocation |
| Repository | Domain | Impl. | Repair | Dev. | Private |
|---|---|---|---|---|---|
| HSTU codebase | Recommendation | ||||
| LitGPT | Language-model training | ||||
| Total | Two ML domains |
| Ordered pair | Repository | Rescue / harm | |||||
|---|---|---|---|---|---|---|---|
| CC CC | HSTU | ||||||
| CC Codex | HSTU | ||||||
| CC Codex | LitGPT | ||||||
| Codex CC | HSTU | ||||||
| Codex Codex | HSTU |
| Process-faithfulness violation | Runs |
|---|---|
| Ran only a subset of selected proposals | 10 |
| Stopped early despite self-recommending continue | 4 |
| Skipped a prescribed user-confirmation gate | 4 |
| Called a tool with off-spec arguments | 3 |
| Redundant re-investigation on loop-back | 3 |
| Merged the codebase and data investigation phases | 2 |
| Injection policy | Early EA | Middle EA | Late EA | CDR | Wrong decision | Input tokens | Cost |
|---|---|---|---|---|---|---|---|
| Current-step only | k | ||||||
| Full protocol | k | ||||||
| Paired difference | k |
| Variant | NDCG@10 | ||
|---|---|---|---|
| Paper baseline | — | 0.2098 | — |
| PRISM+IC stack (ref) | — | 0.2140 | +2.0% |
| Genre FT (leak-free) | 0.05 | 0.2192 | +4.48% |
| Genre FT (leak-free) | 0.10 | 0.2192 | +4.48% |
| Genre FT (leak-free) | 0.20 | 0.2191 | +4.43% |
| Genre from-scratch ∗ | 0.10 | 0.1911 | — |
| Method | NDCG@10 | vs. ref |
|---|---|---|
| Reference (no pref.) | 0.2191 | — |
| DPO | 0.1405 | |
| IPO | 0.2012 | |
| SimPO | 0.1798 |