Organizations: Rutgers University · Independent Researcher · University of California, San Diego · University of Michigan · McGill University · King Fahd University of Petroleum and Minerals
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
Figures & tables
Figure 1: Co-cheating. The training signal and false agreement rise together.
Figure 2: Co-cheating and its mitigation. (a) Incorrect agreement creates false frontier credit. (b) MSV verifies each proposal with source-aware and source-blind samples. (c) CrossFit scores each source group with a solver trained on the other group. The two interventions target label quality and feedback provenance, respectively.
Figure 3: Co-cheating under coupled feedback. (a,b) All 129 steps of label truth TP , solver truth TS , and agreement A ; dashes mark round boundaries. (c,d) False agreement F versus lost credit L : small points are steps, large markers average each round’s 43 step rates, and arrows indicate round order. Rising F with falling L reveals shared-error accumulation. All axes are percentages.
Figure 4: The CrossFit algorithm. Auxiliary solvers train on one source fold and score the other; the main solver trains on all admitted questions.
NQ
TriviaQA
PopQA
HotpotQA
2WikiMQA
MuSiQue
Bamboogle
Average
Qwen3.5-4B
Base
0.380
0.655
0.315
0.350
0.465
0.105
0.416
0.384
Prompting †
0.245
0.525
0.165
0.290
0.460
0.085
0.360
0.304
R1-Instruct †
0.295
0.605
0.195
0.320
0.475
0.105
0.448
0.349
Search-R1 †
0.355
0.650
0.290
0.370
0.510
0.125
0.504
0.401
Dr. Zero
0.390
0.665
0.325
0.365
0.485
0.120
0.448
0.400
Table 1: Downstream search performance after three rounds of self-evolution. Bold denotes the best result and underlining denotes the second-best within each Qwen3.5 block. † Baselines from the Dr. Zero comparison, run on the same Qwen3.5 backbones and evaluated on the same 1,325-question set.
Figure 5: Training dynamics across three rounds. (a,b) Round means of agreement and solver truth; hollow, light, and solid markers denote rounds 1–3. Above the diagonal, agreement is optimistic. (c,d) Faint points retain all 129 steps; solid segments average each round’s 43 step rates, and end labels give round-3 means (%). Gray bands mark proposer phases, short dashes mark round boundaries, and the long dash marks step 62. Full traces and coverage appear in Figures 8 – 9 .
Figure 6: Round-3 audit relative to Dr. Zero. (a,b) Joint truth gains; arrows lead to the combined treatment. (c) Reduction FDr.Zero−F . All values are percentage points.
Figure 7: Mechanism ablations on a fixed replay bank. (a) Truth versus false agreement; arrows compare coupled and source-ID feedback. (b) Accuracy on 3,000 replay questions, mean ± SD over five seeds. Numbers identify methods; shading marks source exclusion.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Configuration
Backbones
Qwen3.5-4B and Qwen3.5-9B
Schedule
Three rounds; 18 proposer and 25 main-solver updates per round
Solver data
1,600 admitted questions per round; five responses per question
Proposer update
64 prompts per update; one selected trajectory per prompt
Sampling
Temperature 0.8 and top- p 0.95 for training trajectories
MSV
Three source-aware and three source-blind samples per proposal
Appendix
Table 2: Training configuration shared across treatments.
H200-hours
Tokens (M)
Judge
Hours
Treatment
4B
9B
Input
Output
req. (k)
4B
9B
Dr. Zero
379
476
696
83
60.4
47
60
MSV
719
903
1,811
195
296.9
90
113
CrossFit (25 total)
515
665
889
107
87.1
64
83
CrossFit (25 per fold)
650
854
1,081
131
113.8
81
107
MSV + CrossFit (25 per fold)
990
1,281
2,196
243
350.3
124
160
Appendix
Table 3: Resource cost per training run. H200-hours are reserved budgets: each run holds eight H200 GPUs for its end-to-end duration, including waiting time, so they exceed the accelerator time actually used; the GPU usage of the external audit service is unknown and not included. Token (millions) and judge-request (thousands) counts are per single-scale run, not summed over the two scales, and exclude training-replay tokens. Hours are the end-to-end wall-clock duration of the full pipeline. “25 total” and “25 per fold” denote the auxiliary-update budget per round.
J/E
TP
TS
A
F
L
Qwen3.5-4B
Dr. Zero
0.859
0.747
0.669
0.710
0.061
0.020
MSV
0.860
0.747
0.708
0.745
0.057
0.020
CrossFit
0.860
0.819
0.686
0.679
0.030
0.038
MSV + CrossFit
0.858
0.843
0.727
0.708
0.020
0.038
Qwen3.5-9B
Appendix
Table 4: Round-3 audit statistics for every treatment: means of the 43 round-3 step rates. J/E is audit coverage; TP , TS , A , F , and L are defined in Section 2 . Figure 6 plots the corresponding changes relative to Dr. Zero, computed from unrounded means; they can therefore differ by 0.1 percentage point from differences of the rounded entries here.
NQ
TriviaQA
PopQA
HotpotQA
2WikiMQA
MuSiQue
Bamboogle
Average
Qwen3.5-4B
Base
0.380
0.655
0.315
0.350
0.465
0.105
0.416
0.384
Dr. Zero Round 1
0.380
0.660
0.325
0.350
0.475
0.125
0.424
0.391
Dr. Zero Round 2
0.395
0.665
0.320
0.360
0.470
0.125
0.440
0.396
Dr. Zero Round 3
0.390
0.665
0.325
0.365
0.485
0.120
0.448
0.400
MSV Round 1
0.385
0.675
0.330
0.355
0.480
0.110
0.432
0.395
Appendix
Table 5: Cover-EM of every round-end main solver; we mark the best performance within each scale in bold. Each cross-fitted treatment shares round 1 with its coupled counterpart.
Qwen3.5-4B
Qwen3.5-9B
Round 1
Round 2
Round 3
Round 1
Round 2
Round 3
Dr. Zero
0.389
0.394
0.397
0.413
0.418
0.423
MSV
0.393
0.399
0.405
0.417
0.423
0.430
CrossFit
0.389
0.436
0.485
0.413
0.462
0.506
MSV + CrossFit
0.393
0.444
0.489
0.417
0.468
0.510
Appendix
Table 6: Micro-averaged Cover-EM of the round-end main solvers; we mark the best performance in bold.
Figure 8: Complete Qwen3.5-4B audit trajectories. Columns identify treatments. Rows show adopted-label truth TP , solver truth TS , and agreement A ; false-agreement mass F ; lost-credit mass L ; and coverage J/E . Every metric is expressed in percent. Shared limits support comparison across treatments and model scales. Gray bands mark proposer phases, short dashes mark round boundaries, and the long dash marks step 62.
Figure 9: Complete Qwen3.5-9B audit trajectories. Layout, metric colors, units, and axis limits match Figure 8 . Every scheduled step is retained. The separate coverage strip prevents audit coverage from obscuring the correctness curves.