Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.
Figures & tables
Figure 1: Recursive self-improvement with multi-agent self-supervision (MASS). The agent, evaluator, and optimizer share the same language model (LM). MASS encodes multi-agent coordination as textual information within a workflow prompt . In the inner loop, the agent generates a workspace under the current workflow prompt, the self-evaluator provides feedback, and the self-optimizer refines the workflow prompt. The outer loop fine-tunes the shared LM on rollouts from the optimized workflow, then restarts the inner loop with the updated model.
Figure 2: Score versus output tokens per task on six public benchmarks. L(0) is the base LM (Qwen3.6-27B); L(1) and L(2) follow one and two cycles of RSI with MASS, all run in the same agent scaffold (qwen-code). Arrows go from L(0) to L(2) , labelled with the ratio of their score per output token: on the four research benchmarks (top row, bottom left), L(2) gains 1.2–1.6 × . Bars show standard deviations over three trials.
Figure 3: Evaluation of RSI with MASS. (a) Cumulative number of tasks for which a workflow is unanimously preferred to the baseline by external strong LLM judges. (b) Task-solving performance across model generations on eight synthetic training tasks, three synthetic test tasks, and two research benchmarks. (c) Evaluator agreement with external strong-model judges.
Executor Lagent
Comparator Lagent
Win rate
L(1)(xt⊕wt(1,i))
L(0)(xt⊕wt(0,i))
0.80
L(0)(xt⊕wt(1,i))
L(0)(xt⊕wt(0,i))
0.64
L(0)(xt⊕wt(1,i))
L(0)(xt)
0.63
L(0)(xt⊕wt(1,i))
L(1)(xt⊕wt(1,i))
0.29
Table 1: Mean win rate over i∈[11] iterations and t∈[12] tasks. wt(g,i) indicates the workflow from MASS generation g in workflow optimization iteration i on task t .
Figure 4: Task-information change from each curve’s value at i=3 : ΔIi=Ii(Zc;T)−I3(Zc;T) . By i=11 , role and instruction information falls, while contract and hop information rises in all four configurations. Vertical scales differ by component.
Figure 6: SFT of L(0) on different datasets; the gray line interpolates the single-agent results.
Figure 7: Workflow-search ablations. (a) Cumulative number of tasks for which L(0)(xt⊕wt(0,i)) is unanimously preferred ( 6/6 external judgments) over L(0)(xt) by iteration i∈[20] . (b) Win rate of L(0)(xt⊕wt(0,i)) against L(0)(xt) .
Table 2: Behavioral evidence of workflow internalization. (a) trained models under bare task prompts. (b) corresponding behaviors in the SFT data.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Split
Study
SFT episodes
51
Train
Macro signals and cross-asset trading
15
53
Train
Trading around monetary-policy announcements
15
56
Train
Yield-curve relative-value strategies
14
202
Train
Protein-interface prediction
15
205
Train
Molecular-dynamics force-field validation
15
302
Train
Learned heuristics for path planning
15
Appendix
Table 3: Task allocation for the main first-cycle student M+ . Task 221 is one of the nine designated training tasks but contributes no SFT data. Each test task is from a different domain.
Stage
Signal and use
Workflow search
Current model’s workspace comparisons guide workflow revisions.
Teacher ranking
Current model’s comparisons determine Bradley–Terry ranks.
Checkpoint selection
Validation loss selects the model checkpoint.
External reporting
GPT-5.5 and Claude Opus 4.8 assess completed workspaces.
Appendix
Table 4: Evaluation roles in homogeneous MASS. The role ablations in Appendix E explicitly change the in-loop evaluator or optimizer.
Task
Split
Pairs
Wins
Losses
Ties
Win share
51
Training
10
53
7
0
88.3%
53
Training
10
52
8
0
86.7%
56
Training
10
38
22
0
63.3%
202
Training
10
44
16
0
73.3%
205
Training
10
12
48
0
20.0%
302
Training
10
45
15
0
75.0%
Appendix
Table 5: Per-task comparison of M+ with the no-harness base. Both models receive the bare task prompt. Each task has ten index-matched rollout pairs, assessed by two reporting judges with three seeds each, yielding 60 verdicts per task. Wins, losses, and ties are counted from the student’s perspective; win share is W/(W+L) . Loop-detection halts remain in the evaluation. Task descriptions are given in Table 3 .
Comparator
Evaluation scope
Pairs
W–L–T
Win share
base
All 11 tasks
110
412–248–0
62.4%
base
8 training tasks
80
315–165–0
65.6%
base
3 test tasks
30
97–83–0
53.9%
Multi-agent student M
5 shared tasks
50
211–88–1
70.6%
Single-agent student S+
5 shared tasks
50
196–104–0
65.3%
Ranked team teachers
All 11 tasks
108
112–536–0
17.3%
Appendix
Table 6: Main results for M+ under the bare task prompt. Each pair receives six reporting verdicts (two judges, three seeds per judge). W–L–T denotes wins, losses, and ties for M+ , and win share excludes ties. The five shared tasks are 51, 53, 56, 307, and 60. The teacher comparison pairs student rollout i with ranked teacher i ; task 207 has only eight eligible teacher references, giving 108 pairs in total. The controls M and S+ are defined in Table 8 .
Benchmark and metric
L(0)
L(1)
L(2)
ScienceAgentBench: success (%)
28.8±1.5
30.4±2.0
31.7±2.0
MLR-Bench: Overall (0–10)
1.58±0.38
1.80±0.30
2.40±0.28
MLR-Bench: papers delivered
31/60
36/60
50/60
AstaBench E2E-Bench-Hard: rubric
0.053±0.002
0.055±0.004
0.067±0.017
AstaBench E2E-Bench-Hard: reports delivered
60/120
69/120
91/120
DSBench: normalized score
0.473±0.029
0.465±0.011
0.482±0.025
Appendix
Table 7: Public-benchmark results for L(k)+qwen-code . Scores are mean ± sample standard deviation over trial means. Task and trial counts are: ScienceAgentBench, 102×3 ; MLR-Bench, 12×5 ; AstaBench E2E-Bench-Hard, 40×3 ; DSBench, 74×3 ; Terminal-Bench 2.0, 89×3 ; and SWE-bench Verified, 500×3 . The last two counts refer to scheduled tasks; scoring denominators and exclusions are given below. Delivery counts pool 60 MLR-Bench or 120 AstaBench executions per model and have no error bar.
Supervised tokens (M)
Win share vs. base (%)
Arm
Episodes
Stored
Processed
All five
Shared train
Task 60
S
17
0.90
1.82
44.0
45.8
36.7
M
17
2.27
3.64
56.0
60.3
39.0
X
34
3.17
8.85
58.8
63.7
40.0
S+
78
4.43
9.11
59.0
61.7
48.3
S++
290
—
42.00
64.0
—
—
Appendix
Table 8: Teacher-composition controls on five shared evaluation tasks. Shared train denotes tasks 51, 53, 56, and 307. Dashes mark quantities not reported here. X has 49 evaluated rollout pairs; the other S,M,S+,M+ comparisons have 50. M+ has broader training-task coverage and a different sampler, so it is not a matched-data control.
Comparison
Tasks
Pairs
W–L
Win rate
All pairs
8
206
701–123
85.1%
Comparable length
6
10
32–8
80.0%
Same workflow
8
100
347–53
86.8%
Same workflow and comparable length
4
4
15–1
93.8%
Appendix
Table 9: Teacher-output comparisons. W–L counts judgments favoring or disfavoring the multi-agent output, with four judgments per pair and no ties. Rows overlap; repeated judgments and reused episodes are not independent experimental replications.
P(flip) by edit-size bin
Configuration
Median edit
ρedit
≤5%
5–15%
15–30%
>30%
Stepwise LLM Feedback: Leval(wi,wi+1)
All base
16%
−0.12
0.22 (9)
0.32 (34)
0.23 (31)
0.32 (22)
GPT-5.5 evaluator
15%
+0.03
0.23 (13)
0.32 (47)
0.24 (34)
0.31 (26)
GPT-5.5 evaluator and optimizer
39%
−0.05
— (0)
— (0)
0.40 (30)
0.31 (78)
Baseline LLM Feedback: Leval(wi,w∅)
Appendix
Table 10: Workflow edit size and changes in external judgments. Edit size is the fraction of tokens rewritten under longest-common-block alignment. ρedit is its Spearman correlation with the change in wins out of six. A flip changes ≥5/6 wins to ≤1/6 , or the reverse. Parentheses give the number of steps in each bin. The last row also uses retained-workflow and design-difference history.
Task
Iteration
Execution outcome
W–L–T vs. base
53
16
Completed
6–0–0
53
17
Loop halt: repeated identical shell-tool calls
0–6–0
Appendix
Table 11: Two executions of identical workflow and rendered prompt texts. The six verdicts in each row assess one workspace against the same bare-agent reference. The later failure is a recorded agent loop halt.
Figure 8: Retained best-workflow version by task. A plateau keeps the incumbent; an advance records its replacement. The information analysis samples this path rather than every proposed workflow.
Unadjusted
Permutation-adjusted
Component
i=3
i=6
i=9
i=11
i=3
i=6
i=9
i=11
text-embedding-3-large
role
1.26
1.17
1.11
1.09↓
0.42
0.43
0.40
0.38↓
instruction
2.56
2.40
2.37
2.36↓
1.74
1.64
1.62
1.61↓
contract
2.99
2.93
3.06
3.29↑
2.24
2.18
2.27
2.50↑
hop
5.23
6.00
6.22
6.22↑
4.72
5.34
5.54
5.54↑
Appendix
Table 12: Task-information estimates I(Zc;T) (nats) along the crown path. In all four encoder/estimator configurations, roles and instructions decrease while contracts and hops increase between i=3 and i=11 (arrows). Unadjusted estimates use the full sample; permutation-adjusted estimates subtract a shuffled baseline using 43-agent samples (Appendix F ).
Unadjusted
Permutation-adjusted
Statistic
i=3
i=6
i=9
i=11
i=3
i=6
i=9
i=11
text-embedding-3-large
TC(Z)
11.82
12.18
11.96
11.62↓
5.76
5.73
5.70
5.61↓
TC(Z∣T)
5.51
5.45
5.33
4.95↓
3.11
3.09
3.06
2.83↓
all-mpnet-base-v2
TC(Z)
10.34
10.72
10.36
9.97↓
4.46
4.49
4.38
4.16↓
Appendix
Table 13: Component redundancy (nats), with the same crown workflows and estimator variants as Table 12 . From i=3 to i=11 , unconditional TC(Z) decreases in all four configurations and task-centered TC(Z∣T) in three; the permutation-adjusted MPNet estimate increases.
Model
Prompt
Episodes
D
H
Revision
Main edit
Base
Bare
110
1
1
0
1
M+
Bare
110
48
42
10
14
Base
Generic directive
110
110
103
40
32
M+
Generic directive
110
110
106
56
13
Base
Teacher workflow + directive
114
114
98
57
19
Appendix
Table 14: Episodes with each recorded behavior. Counts use all episodes in the row. Directive comparisons match tasks and generation seeds; teachers are selected episodes on eight training tasks, so their row is descriptive rather than a controlled comparison.
Figure 9: Observed artifact dependencies in bare-prompt test task 305 (seed 2310005). Each arrow requires a producer write, producer return, and consumer read. Numbers indicate launch order; branches indicate dependencies, not parallel execution. BC is behavior cloning; PDDL is Planning Domain Definition Language.
Method
Update and feedback
Relationship to MASS
Self-Rewarding LMs ( Yuan et al., 2024 )
Iterative preference training using the model’s own response scores.
Establishes joint improvement of generation and evaluation, without workflow search.
Multi-Agent Evolve ( Chen et al., 2025a )
A shared backbone learns question generation, solving, and judging through reinforcement learning.
Establishes multi-agent learning with a shared self-evaluator; its proposer generates questions, rather than execution workflows.
ADAS / AFlow / GPTSwarm ( Hu et al., 2025 ; Zhang et al., 2025 ; Zhuge et al., 2024 )
Search over agent programs, workflow graphs, or prompts and communication edges.
Establishes automatic design of agentic computation; MASS learns shared weights from the resulting executions.
RHI ( Lee et al., 2026 )
Prompt-level workflow revisions guided by pairwise comparisons of task artifacts.
Supplies the workflow-search method and information-flow motivation used here; MASS adds trace-based post-training.
TTHE ( Nie et al., 2026 )
One frozen LLM solves tasks, proposes executable harness edits, and judges them using execution-derived proxies.
A Claude Sonnet feedback agent edits the scaffold and initiates weight updates to a gpt-oss target, using task verifiers.
Establishes harness–weight co-updating with a separate improvement model and verifier.
Appendix
Table 15: Selected related methods, compared by their reported mechanisms. The rows describe methods, not matched experimental baselines. For MASS, the homogeneous condition applies to the in-loop executor, evaluator, and optimizer; model-role ablations vary this assignment.
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.
Yaxin Du, Xiyuan Yang, Zhifan Zhou +10
Shanghai Jiao Tong University · Carnegie Mellon University · University of Waterloo
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
Peng Xia, Rujun Han, Zifeng Wang +11
Google Cloud AI Research · UNC-Chapel Hill · Stanford University +1
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.