Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.
Figures & tables
Figure 1: Recursive self-improvement with multi-agent self-supervision (MASS). The agent, evaluator, and optimizer share the same language model (LM). MASS encodes multi-agent coordination as textual information within a workflow prompt . In the inner loop, the agent generates a workspace under the current workflow prompt, the self-evaluator provides feedback, and the self-optimizer refines the workflow prompt. The outer loop fine-tunes the shared LM on rollouts from the optimized workflow, then restarts the inner loop with the updated model.
Figure 2: Score versus output tokens per task on six public benchmarks. L(0) is the base LM (Qwen3.6-27B); L(1) and L(2) follow one and two cycles of RSI with MASS, all run in the same agent scaffold (qwen-code). Arrows go from L(0) to L(2) , labelled with the ratio of their score per output token: on the four research benchmarks (top row, bottom left), L(2) gains 1.2–1.6 × . Bars show standard deviations over three trials.
Figure 3: Evaluation of RSI with MASS. (a) Cumulative number of tasks for which a workflow is unanimously preferred to the baseline by external strong LLM judges. (b) Task-solving performance across model generations on eight synthetic training tasks, three synthetic test tasks, and two research benchmarks. (c) Evaluator agreement with external strong-model judges.
Executor Lagent
Comparator Lagent
Win rate
L(1)(xt⊕wt(1,i))
L(0)(xt⊕wt(0,i))
0.80
L(0)(xt⊕wt(1,i))
L(0)(xt⊕wt(0,i))
0.64
L(0)(xt⊕wt(1,i))
L(0)(xt)
0.63
L(0)(xt⊕wt(1,i))
L(1)(xt⊕wt(1,i))
0.29
Table 1: Mean win rate over i∈[11] iterations and t∈[12] tasks. wt(g,i) indicates the workflow from MASS generation g in workflow optimization iteration i on task t .
Figure 4: Task-information change from each curve’s value at i=3 : ΔIi=Ii(Zc;T)−I3(Zc;T) . By i=11 , role and instruction information falls, while contract and hop information rises in all four configurations. Vertical scales differ by component.
Figure 6: SFT of L(0) on different datasets; the gray line interpolates the single-agent results.
Figure 7: Workflow-search ablations. (a) Cumulative number of tasks for which L(0)(xt⊕wt(0,i)) is unanimously preferred ( 6/6 external judgments) over L(0)(xt) by iteration i∈[20] . (b) Win rate of L(0)(xt⊕wt(0,i)) against L(0)(xt) .
Table 2: Behavioral evidence of workflow internalization. (a) trained models under bare task prompts. (b) corresponding behaviors in the SFT data.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Split
Study
SFT episodes
51
Train
Macro signals and cross-asset trading
15
53
Train
Trading around monetary-policy announcements
15
56
Train
Yield-curve relative-value strategies
14
202
Train
Protein-interface prediction
15
205
Train
Molecular-dynamics force-field validation
15
302
Train
Learned heuristics for path planning
15
Appendix
Table 3: Task allocation for the main first-cycle student M+ . Task 221 is one of the nine designated training tasks but contributes no SFT data. Each test task is from a different domain.
Stage
Signal and use
Workflow search
Current model’s workspace comparisons guide workflow revisions.
Teacher ranking
Current model’s comparisons determine Bradley–Terry ranks.
Checkpoint selection
Validation loss selects the model checkpoint.
External reporting
GPT-5.5 and Claude Opus 4.8 assess completed workspaces.
Appendix
Table 4: Evaluation roles in homogeneous MASS. The role ablations in Appendix E explicitly change the in-loop evaluator or optimizer.
Task
Split
Pairs
Wins
Losses
Ties
Win share
51
Training
10
53
7
0
88.3%
53
Training
10
52
8
0
86.7%
56
Training
10
38
22
0
63.3%
202
Training
10
44
16
0
73.3%
205
Training
10
12
48
0
20.0%
302
Training
10
45
15
0
75.0%
Appendix
Table 5: Per-task comparison of M+ with the no-harness base. Both models receive the bare task prompt. Each task has ten index-matched rollout pairs, assessed by two reporting judges with three seeds each, yielding 60 verdicts per task. Wins, losses, and ties are counted from the student’s perspective; win share is W/(W+L) . Loop-detection halts remain in the evaluation. Task descriptions are given in Table 3 .
Comparator
Evaluation scope
Pairs
W–L–T
Win share
base
All 11 tasks
110
412–248–0
62.4%
base
8 training tasks
80
315–165–0
65.6%
base
3 test tasks
30
97–83–0
53.9%
Multi-agent student M
5 shared tasks
50
211–88–1
70.6%
Single-agent student S+
5 shared tasks
50
196–104–0
65.3%
Ranked team teachers
All 11 tasks
108
112–536–0
17.3%
Appendix
Table 6: Main results for M+ under the bare task prompt. Each pair receives six reporting verdicts (two judges, three seeds per judge). W–L–T denotes wins, losses, and ties for M+ , and win share excludes ties. The five shared tasks are 51, 53, 56, 307, and 60. The teacher comparison pairs student rollout i with ranked teacher i ; task 207 has only eight eligible teacher references, giving 108 pairs in total. The controls M and S+ are defined in Table 8 .
Benchmark and metric
L(0)
L(1)
L(2)
ScienceAgentBench: success (%)
28.8±1.5
30.4±2.0
31.7±2.0
MLR-Bench: Overall (0–10)
1.58±0.38
1.80±0.30
2.40±0.28
MLR-Bench: papers delivered
31/60
36/60
50/60
AstaBench E2E-Bench-Hard: rubric
0.053±0.002
0.055±0.004
0.067±0.017
AstaBench E2E-Bench-Hard: reports delivered
60/120
69/120
91/120
DSBench: normalized score
0.473±0.029
0.465±0.011
0.482±0.025
Appendix
Table 7: Public-benchmark results for L(k)+qwen-code . Scores are mean ± sample standard deviation over trial means. Task and trial counts are: ScienceAgentBench, 102×3 ; MLR-Bench, 12×5 ; AstaBench E2E-Bench-Hard, 40×3 ; DSBench, 74×3 ; Terminal-Bench 2.0, 89×3 ; and SWE-bench Verified, 500×3 . The last two counts refer to scheduled tasks; scoring denominators and exclusions are given below. Delivery counts pool 60 MLR-Bench or 120 AstaBench executions per model and have no error bar.
Supervised tokens (M)
Win share vs. base (%)
Arm
Episodes
Stored
Processed
All five
Shared train
Task 60
S
17
0.90
1.82
44.0
45.8
36.7
M
17
2.27
3.64
56.0
60.3
39.0
X
34
3.17
8.85
58.8
63.7
40.0
S+
78
4.43
9.11
59.0
61.7
48.3
S++
290
—
42.00
64.0
—
—
Appendix
Table 8: Teacher-composition controls on five shared evaluation tasks. Shared train denotes tasks 51, 53, 56, and 307. Dashes mark quantities not reported here. X has 49 evaluated rollout pairs; the other S,M,S+,M+ comparisons have 50. M+ has broader training-task coverage and a different sampler, so it is not a matched-data control.
Comparison
Tasks
Pairs
W–L
Win rate
All pairs
8
206
701–123
85.1%
Comparable length
6
10
32–8
80.0%
Same workflow
8
100
347–53
86.8%
Same workflow and comparable length
4
4
15–1
93.8%
Appendix
Table 9: Teacher-output comparisons. W–L counts judgments favoring or disfavoring the multi-agent output, with four judgments per pair and no ties. Rows overlap; repeated judgments and reused episodes are not independent experimental replications.
P(flip) by edit-size bin
Configuration
Median edit
ρedit
≤5%
5–15%
15–30%
>30%
Stepwise LLM Feedback: Leval(wi,wi+1)
All base
16%
−0.12
0.22 (9)
0.32 (34)
0.23 (31)
0.32 (22)
GPT-5.5 evaluator
15%
+0.03
0.23 (13)
0.32 (47)
0.24 (34)
0.31 (26)
GPT-5.5 evaluator and optimizer
39%
−0.05
— (0)
— (0)
0.40 (30)
0.31 (78)
Baseline LLM Feedback: Leval(wi,w∅)
Appendix
Table 10: Workflow edit size and changes in external judgments. Edit size is the fraction of tokens rewritten under longest-common-block alignment. ρedit is its Spearman correlation with the change in wins out of six. A flip changes ≥5/6 wins to ≤1/6 , or the reverse. Parentheses give the number of steps in each bin. The last row also uses retained-workflow and design-difference history.
Task
Iteration
Execution outcome
W–L–T vs. base
53
16
Completed
6–0–0
53
17
Loop halt: repeated identical shell-tool calls
0–6–0
Appendix
Table 11: Two executions of identical workflow and rendered prompt texts. The six verdicts in each row assess one workspace against the same bare-agent reference. The later failure is a recorded agent loop halt.
Figure 8: Retained best-workflow version by task. A plateau keeps the incumbent; an advance records its replacement. The information analysis samples this path rather than every proposed workflow.
Unadjusted
Permutation-adjusted
Component
i=3
i=6
i=9
i=11
i=3
i=6
i=9
i=11
text-embedding-3-large
role
1.26
1.17
1.11
1.09↓
0.42
0.43
0.40
0.38↓
instruction
2.56
2.40
2.37
2.36↓
1.74
1.64
1.62
1.61↓
contract
2.99
2.93
3.06
3.29↑
2.24
2.18
2.27
2.50↑
hop
5.23
6.00
6.22
6.22↑
4.72
5.34
5.54
5.54↑
Appendix
Table 12: Task-information estimates I(Zc;T) (nats) along the crown path. In all four encoder/estimator configurations, roles and instructions decrease while contracts and hops increase between i=3 and i=11 (arrows). Unadjusted estimates use the full sample; permutation-adjusted estimates subtract a shuffled baseline using 43-agent samples (Appendix F ).
Unadjusted
Permutation-adjusted
Statistic
i=3
i=6
i=9
i=11
i=3
i=6
i=9
i=11
text-embedding-3-large
TC(Z)
11.82
12.18
11.96
11.62↓
5.76
5.73
5.70
5.61↓
TC(Z∣T)
5.51
5.45
5.33
4.95↓
3.11
3.09
3.06
2.83↓
all-mpnet-base-v2
TC(Z)
10.34
10.72
10.36
9.97↓
4.46
4.49
4.38
4.16↓
Appendix
Table 13: Component redundancy (nats), with the same crown workflows and estimator variants as Table 12 . From i=3 to i=11 , unconditional TC(Z) decreases in all four configurations and task-centered TC(Z∣T) in three; the permutation-adjusted MPNet estimate increases.
Model
Prompt
Episodes
D
H
Revision
Main edit
Base
Bare
110
1
1
0
1
M+
Bare
110
48
42
10
14
Base
Generic directive
110
110
103
40
32
M+
Generic directive
110
110
106
56
13
Base
Teacher workflow + directive
114
114
98
57
19
Appendix
Table 14: Episodes with each recorded behavior. Counts use all episodes in the row. Directive comparisons match tasks and generation seeds; teachers are selected episodes on eight training tasks, so their row is descriptive rather than a controlled comparison.
Figure 9: Observed artifact dependencies in bare-prompt test task 305 (seed 2310005). Each arrow requires a producer write, producer return, and consumer read. Numbers indicate launch order; branches indicate dependencies, not parallel execution. BC is behavior cloning; PDDL is Planning Domain Definition Language.
Method
Update and feedback
Relationship to MASS
Self-Rewarding LMs ( Yuan et al., 2024 )
Iterative preference training using the model’s own response scores.
Establishes joint improvement of generation and evaluation, without workflow search.
Multi-Agent Evolve ( Chen et al., 2025a )
A shared backbone learns question generation, solving, and judging through reinforcement learning.
Establishes multi-agent learning with a shared self-evaluator; its proposer generates questions, rather than execution workflows.
ADAS / AFlow / GPTSwarm ( Hu et al., 2025 ; Zhang et al., 2025 ; Zhuge et al., 2024 )
Search over agent programs, workflow graphs, or prompts and communication edges.
Establishes automatic design of agentic computation; MASS learns shared weights from the resulting executions.
RHI ( Lee et al., 2026 )
Prompt-level workflow revisions guided by pairwise comparisons of task artifacts.
Supplies the workflow-search method and information-flow motivation used here; MASS adds trace-based post-training.
TTHE ( Nie et al., 2026 )
One frozen LLM solves tasks, proposes executable harness edits, and judges them using execution-derived proxies.
A Claude Sonnet feedback agent edits the scaffold and initiates weight updates to a gpt-oss target, using task verifiers.
Establishes harness–weight co-updating with a separate improvement model and verifier.
Appendix
Table 15: Selected related methods, compared by their reported mechanisms. The rows describe methods, not matched experimental baselines. For MASS, the homogeneous condition applies to the in-loop executor, evaluator, and optimizer; model-role ablations vary this assignment.