Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.
Figures & tables
Figure 2 : Overview of RSI-Master . The Experiment OS defines permitted operations for data development, training, and evaluation and maintains persistent, linked experimental records. Reviewer-Guided Research Orchestration grows a DAG of research directions: Workers explore directions through experiments, Reviewers compare the resulting evidence, and the Main Agent uses their reviews to continue, verify, branch, or revise directions.
Category
Num
Representative tool
Function
Data management
6
worker_add_data
Register and retrieve data artifacts.
Execution & state
13
state_sync_results
Manage execution and synchronize state.
Research & evidence
8
compare_eval_samples
Retrieve context and compare evidence.
Review management
4
inspect_reviewer_report
Store and retrieve experiment reviews.
External discovery
22
hf_search_datasets
Discover datasets and external information.
Table 1 : Experiment OS tool categories in the current implementation. Counts include optional tools and are deduplicated across agent roles. External connectors are counted separately from core tools.
Table 2: Performance of autonomous post-training systems on PostTrainBench, starting from Qwen3-4B-Base. Qwen3-4B-Base and Qwen3-4B-Instruct are included as references. Avg. is the unweighted mean of the seven task scores. Higher is better; the best results among autonomous systems are highlighted in bold , including ties.
Figure 3 : Comparison of RSI-Master and the human-developed Qwen3.5-35B-A3B Instruct reference on frontier benchmarks. Hatched regions indicate available Base score and Avg. is the unweighted mean.
Figure 4 : Cross-domain post-training results. RSI-Master is trained on Qwen3-4B-Base; Base denotes the initial Qwen3-4B-Base checkpoint, and Qwen3-4B (Instruct) is the instruct reference. RSI-Master improves over Base on all 13 completed tasks and exceeds Instruct on seven, supporting its applicability across domains. Avg is the unweighted mean over these 13 tasks.
Configuration
ExpOS
Worker
Reviewer
AIME 2025
Arena Hard
BFCL
GPQA Main
GSM8K
HealthBench
HumanEval
Avg ↑
w/o ExpOS
×
✓
✓
23.33
22.31
62.29
37.95
88.70
26.98
65.24
46.69
w/o Worker
✓
×
✓
10.00
48.15
61.46
33.93
78.92
14.17
67.68
44.90
w/o Reviewer
✓
✓
×
13.33
24.66
63.73
39.96
91.58
16.05
63.41
44.67
RSI-Master
✓
✓
✓
23.33
49.95
64.50
41.29
92.20
35.79
74.39
54.49
Table 3: Component ablation of RSI-Master on PostTrainBench. ✓ / × indicates whether ExpOS, Worker, and Reviewer are enabled. Avg is the mean performance across all 7 tasks.
Figure 5 : Experimental integrity and research behavior. (a) Protocol-sensitive behavior counts across five agent harnesses, grouped by training, inference, evaluation, and reporting; the curve shows each stage’s share of recorded occurrences. (b) RSI-Master tool-use counts across five progress intervals, grouped by behavior type; red markers indicate episodes in which Reviewer feedback changed the subsequent action.
Method
ExpOS
AIME 2025
Arena Hard
BFCL
GPQA
GSM8K
HealthBench
HumanEval
Avg ↑
Hacking Rate ↓
Claude Code
×
8.89
40.10
38.74
36.16
88.80
17.45
75.61
43.68
8.2%
Claude Code
✓
17.78
43.12
64.91
35.71
92.60
40.62
64.03
51.25
4.7%
Codex
×
16.67
21.44
59.63
35.71
82.40
21.72
65.24
43.26
32.9%
Codex
✓
3.33
69.49
51.12
35.27
84.20
21.03
40.24
43.53
10.5%
RSI-Master
✓
23.33
49.95
64.50
41.29
92.20
35.79
74.39
54.49
0.0%
Table 4 : Comparison of agent frameworks with and without ExpOS on PostTrainBench, with RSI-Master included as a reference. Avg. denotes the macro average across all seven tasks. Bold indicates the best score in each task column and the lowest hacking rate.
Figure 6 : Research-budget scaling. (a) Time-scaling plot shows mean best-so-far score overtime. (b) Recorded mean score gains from 6 to 12 hours across PostTrainBench.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Domain
Evaluation Method
LEXam [ 13 ]
Law
Multiple-choice accuracy (16 options).
LegalBench [ 14 ]
Law
Sample-weighted balanced accuracy over five tasks, using an LLM judge under the DataPrep-Bench protocol.
MedCaseReasoning [ 44 ]
Medicine
LLM-judged diagnostic equivalence accuracy; clinical reasoning-point recall is recorded separately.
MedXpertQA [ 59 ]
Medicine
Multiple-choice accuracy (five options).
WMDP-Bio [ 23 ]
Biology
Multiple-choice accuracy (four options).
EconLogicQA [ 35 ]
Economics
Exact match of the predicted event-ordering sequence.
Appendix
Table 5: Domain-specific benchmarks and evaluation methods for Qwen3-4B. Each task is optimized in a separate run. The methods describe our evaluation protocols; CKQA denotes FinCDM-CPA-KQA.
Figure 7 : Composition of recorded ExpOS tool usage across six functional categories.
Tool name(s)
Function
Data management (6)
worker_show_datapool , reviewer_show_datapool
List shared data-pool entries.
worker_inspect_data , reviewer_inspect_data
Inspect a registered data entry.
worker_add_data
Register metadata and optionally copy a data artifact, recording its hash.
worker_update_readme
Append descriptive or provenance information to a data entry.
Execution and state (13)
Appendix
Table 6 : Tool inventory presented in this paper, including optional interfaces. Grouped names denote separate tools with related functions.
Figure 8 : Human intervention during autonomous research on AIME 2025. Curves show the best score attained over wall-clock time. Expert guidance enters at the Main Agent’s planning stage at exp_0006 . The guided continuation improves to 26.67% within 1.65 hours, while the original Full run plateaus at 23.33% through 12 hours.
Figure 9 : Evaluation-guided parameter injection during a selected stage of a Kimi Swarm HealthBench run.
Figure 10 : A selected stage of RSI-Master’s autonomous BFCL optimization.